MMDF-Net: A Multimodal Framework for Real-Time Deepfake Detection Using Visual and Audio Features

Muhammad Yaqoob Javed, Fida Hussain Dahri, Asif Ali Laghari, Jameel Ahmed Bhutto, Sajjad Hussain Bhutto, Nisar Ahmed Dahri · 2025

The rapid evolution of deepfake technology necessitates detection frameworks capable of leveraging diverse modalities to ensure robust and real-time performance. This research introduces MMDF-Net, a novel multimodal framework for deepfake detection that integrates visual and audio features. MMDF-Net employs MobileNetV2 for lightweight local facial feature extraction and the Swin Transformer to capture global facial relationships, ensuring comprehensive visual representation. The Wav2vec model is used for audio feature extraction to balance visual data, enabling a holistic approach to detection and using the multimodal fusion mechanism to integrate these features, facilitating an accurate and reliable real-time deepfake framework. The proposed MMDF-Net framework achieves impressive performance on the LAV-DF and TVIL datasets, with LAV-DF yielding an accuracy of 97.0% and TVIL achieving 95.5%. Both datasets demonstrate robust detection capabilities, with the model maintaining high precision, recall, and F1-score while processing samples in less than 100ms, explaining the significance of combining visual and audio modalities for efficient and effective deepfake detection.

Read the paper · More papers on PaperTik