Vision-based Multimodal Deepfake Detection using Explainable AI

Justin Joseph, Sujitha Juliet, J. Anitha · 2025

The rise of deepfake technology has raised critical issues concerning digital security, disinformation, and identity theft. Detection models of deepfakes have conventionally relied greatly on convolutional neural networks (CNNs), powerful as they are, but they lack transparency and suffer from generalizability to multiple manipulation techniques. This paper proposes a vision-based multimodal deepfake detection architecture that uses CNNs, Vision Transformers (ViTs), and Explainable AI (XAI) techniques to enhance performance as well as trustworthiness. The proposed model uses spatiotemporal inconsistencies in facial texture, eye blinking, and visual consistency for enhanced detection. The model's mean accuracy of 97.2%, as demonstrated by experimental findings on the FaceForensics++, Celeb-DF, and DFDC datasets, is superior to the most advanced deepfake detection techniques. In addition, Grad-CAM and SHAP-based visual explanations have explainable results, further increasing forensic usability. This research advances towards more accurate, explainable, and effective detection against evolving adversarial attacks and defending digital integrity and fighting misinformation will help achieve Sustainable Development Goal (SDG) 16: Peace, Justice, and Strong Institutions.

Read the paper · More papers on PaperTik