Deepfake Video Prediction Using Attention-Based CNN and Mel-Frequency Cepstral Coefficients
S Geerthik, Senthil G. A, D. Jayashree, J Abinaya · 2024
Deepfake technology seriously threatens the integrity of digital media since it makes it harder to detect fraudulent deepfake videos using traditional methods. In this research, we describe a unique method for audio and visual analysis utilizing Mel-Frequency Cepstral Coefficients (MFCCs) and Attention-based CNN to detect fraudulent deepfakes. Our methodology addresses the evolving field of deepfake manipulation and offers a more comprehensive and dependable detection framework by combining audio and visual modalities. Utilizing the FakeAVceleb dataset, a carefully chosen collection of audio-visual deepfake videos created especially for research is one of the keystones of our methodology. The breadth of actual and deepfake sounds and videos provided by the FakeAVceleb dataset makes it possible to test and evaluate our detection model in great detail. We use attention-based CNN to analyze extracted visual features from video frames, with a focus on regions that can be manipulated to detect subtle audio features that indicate deepfake manipulation, thereby increasing our model's overall discriminative capability. We assess how well our method works for identifying deepfake audio and video. Our results demonstrate the efficacy of combining audio and visual analysis in deepfake detection, as our model achieves promising performance metrics including an accuracy of 94.00% using the FakeAVceleb dataset.