The Deepfake Dilemma: Enhancing Deepfake Detection with Vision Transformers
D. Sumathi, Ashu Singh, Arpita Sinha, Dhea Nerizza Aditya, Mohammed Riyaan K F · 2025
The emergence of deepfake videos at an alarming pace has compromised the integrity of digital multimedia and mandates progressive research into detection strategies. A new forensic method for subjecting face tampering detection in videos using the FaceForensics++ dataset is introduced in this work. A Convolutional Neural Network (CNN) based architecture is enhanced with Vision Transformer encoders for identifying facial manipulations introduced at microscopic precision. The hybrid model integrates CNNs and ViTs, allowing for the capturing of both local and global features that help in detecting subtle manipulations. ViTs have a built-in self-attention mechanism, which allows the model to concentrate on essential facial characteristics, thereby enhancing precision even if the manipulations are confirmed in less than optimal situations. Additionally, the pre-trained Gemini 1.5 Pro model is fine-tuned for optimal performance. The classification of faces into authentic or fake is accurately done using the key points by 90%accuracy as shown through experimentation. Moreover, the Haar cascade method is used for human face detection technique as part of integrated design which augments its complete application in real-world functionalities.