Deep Learning-Based Detection of AI-Generated Voices Using Spectral Features
Mohammad Shike, Muhammad Irfan, Suk Jin Lee, Ahmad Salman · 2025
This paper presents a state-of-the-art deep learning model for detecting AI-generated voices using spectrogram-based classification. We implement a transfer learning strategy based on the EfficientNetB0 architecture to analyze log-mel spectrograms derived from audio samples. Our methodology leverages a comprehensively curated dataset that integrates mul-tiple public sources, totaling 35,284 audio files equally distributed between authentic and synthetic voices. Notably, this dataset is significantly larger, well-balanced, and more diverse than existing voice datasets, providing a robust foundation for training and evaluation. Through a two-phase training approach, our model achieves a 99.40% validation accuracy, surpassing existing approaches in synthetic voice detection, such as those using speech pause patterns. This higher validation accuracy under-scores the effectiveness of our method in accurately distinguishing between authentic and synthetic voices. The model demonstrates exceptional robustness with balanced precision and recall scores of 99.30% and 99.49%, respectively. Additional experiments with VGG16, VGG19, and ResNet50 further validate our approach. This work addresses the critical need for reliable detection methods of AI-generated voices amid the rising prevalence of synthetic media.