A Comprehensive Approach to Deepfake Audio Detection: Using Feature Fusion and Deep Learning
Sharmin Akter Momu, Rafiur Rahman Siddiqui, Shakib Sadat Shanto, Zishan Ahmed · 2024
The rise of deepfake audio has increased concerns regarding the authenticity and integrity of the audio that we hear now a day. Our research proposes a multi-feature fusion approach detecting deepfake audio by integrating Mel-frequency cepstral coefficients (MFCCs), and linear frequency cepstral coefficients (LFCCs), with a deep learning architecture. The model makes the most of VGG16’s robust image-based feature extraction capabilities by transforming acoustic signals into visual spectrogram representations. Both MFCCs and LFCCs features are processed through Temporal Convolutional Networks (TCNs) to capture temporal dependencies, after which the outputs are concatenated for analysis. To understand both the short-term and long-term temporal patterns in the audio data, we used a Bidirectional Long Short-Term Memory (BiLSTM) Model. Our proposed method accomplished a significant test accuracy of 98.87% and a lower Equal Error Rate (EER) of 1.79% for In-The-Wild audio deepfake dataset. On the other hand, for the ASVspoof 2019 dataset, we achieved an accuracy of 97.21% and EER of 2.69%. The results show our approach of combining multiple audio features and utilizing deep learning architecture results in enhanced accuracy and reliable deepfake audio detection.