Towards Enhanced Deep Fake Detection: Integrating Vision Transformers and EfficientNet Deep Features
S. Arun Kumar, S. Thahir Irfan, S. Sanjaykumar, S. L. Shagar Shree Raja, S. Sasikala · 2025
With the fast growth of artificial intelligence development lately, it has become a significant challenge to differentiate between authentic and manipulated digital content. Deep fakes refer to complex algorithms that create realistic fake images, videos, or audio. They pose threats to information integrity, cybersecurity, and public trust. This work addresses the growing need for robust detection mechanisms to combat deep fake technologies. A system for deep fake detection for image is proposed in this work. For deep fake image detection, Vision Transformer (ViT) model is used for global feature extraction and EfficientNet model is used for local feature extraction. The results are then combined and used for classification process. Further, a comparative study is presented for different types of classifiers. Among the various classifiers implemented, Sigmoid SVM showed the best accuracy (85.97%). The models have been designed using public datasets. Metrics such as Accuracy, Loss, Precision, Sensitivity (recall), Specificity, F1-score, MCC, BCR and Cohen-Kappa are evaluated for the proposed system.