PViT: A Hybrid Model for Deepfake Face Detection using Patch Vision Transformers and Deep Learning

Iqra Ambreen, Muhammad Aatif, Zunera Jalil, Farkhund Iqbal, Andrew Marrington · 2025

The proliferation of AI-generated deepfakes, particularly facial image forgeries, poses a significant threat to digital security by facilitating misinformation, identity theft, and privacy breaches. Traditional detection approaches, primarily based on Convolutional Neural Networks (CNNs), often exhibit limited effectiveness when confronted with highly refined or subtle manipulations, leading to compromised detection performance. To address this challenge, this study explores the application of Vision Transformers (ViTs), which leverage selfattention mechanisms to capture fine-grained inconsistencies in visual patterns. This research proposed a hybrid deepfake detection model that integrates patch-oriented ViTs with CNN architectures to improve discriminative feature extraction. Experimental evaluation on benchmark datasets demonstrates that the proposed model achieved a detection accuracy 99%, precision $\mathbf{9 9 \%}$, recall $\mathbf{9 9\%}$, F1-Score $\mathbf{9 9\%}$ on a validation set comprising 76,161 facial images, outperforming conventional CNN-based methods. These results highlight the potential of transformerbased architectures in advancing the robustness and reliability of deepfake detection systems, thereby contributing to the protection of digital authenticity and information integrity.

Read the paper · More papers on PaperTik