Deep Fake Audio Detection Framework Using MFCCs, Chroma Features, and Spectrogram Images
Nora Bakken, S. Birendra Singh, Makwana Prashant, Tapadhir Das · 2025
Deepfake audio detection is critical for combating misinformation in a digital age overrun with dubious media, especially in the new age of AI. This study investigates the performance of different machine learning approaches for classifying deepfake audio, leveraging both Mel-spectrogram representations and Mel-frequency Cepstral Coefficients as audio features. We utilized ResNet50 and VGG16 architectures to classify Mel spectrogram images and applied gradient boosting and support vector machines (SVM) on MFCCs. To SVM and Gradient Boosting, we additionally supplied chroma features, which yielded excellent results. While ResNet50 and VGG16 achieved high accuracy for original audio spectrograms, their performance degraded significantly on 2-second clips and rerecorded samples. In contrast, gradient boosting on MFCCs with chroma features achieved over 95% accuracy, and SVM on MFCCs with chroma features yielded 99% accuracy across original, normalized, 2second, and rerecorded datasets. These insights contribute to the development of reliable detection systems for real-world scenarios.