Multimodal Deepfake Detection using Deep-Convolutional Neural Networks and Mel-Frequency Cepstral Coefficients

Noel Martin, Sharun Raj Nambayil, U Devikrishna, Joel Jismon P, Fasila K.A · 2024

Deepfake technology has raised significant concerns regarding the authenticity of multimedia content, posing threats to security, privacy, and trust in digital media. This work focuses on the development of robust and efficient deepfake detection algorithms for both audio and video content, leveraging the power of Deep Learning. For video, our approach uses Deep Convolutional Neural Networks (DCNNs) to analyze facial features, speech patterns, and contextual information, enabling the identification of manipulated videos, achieving an accuracy of 95.9%. For audio, we utilize Mel-Frequency Cepstral Coefficients (MFCCs) feature extraction, traditional Convolutional Neural Network (CNN) along with LSTM, and Spectrogram analysis to detect inconsistencies in vocal patterns, emphasizing both textual and acoustic features, reaching an accuracy of 98.2%. Then we implement a Multimodal classification method which consists of a combination of three neural networks: one which extracts features from the visual modality, another which extracts features from the textual modality, and a final network which decides which one of the two modalities is more informative, and then performs classification by paying attention to the more informative modality. The proposed methods demonstrate promising results in combating the rising threat of deepfake manipulation in multimedia content, contributing to the preservation of trust and integrity in the digital age.

Read the paper · More papers on PaperTik