Hybrid CNN-Transformer Architecture for Robust Deepfake Detection: A Keyframe-Based Evaluation

S Manasa · INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2025

Abstract - The proliferation of Deepfake content presents a significant threat to digital integrity and media authenticity. To address this challenge, we present a comprehensive evaluation of four deep learning architectures—Convolutional Neural Networks (CNN), Transformer-based models, CNN integrated with Long Short-Term Memory (CNN+LSTM), and a novel hybrid CNN–Transformer model—specifically applied to Deepfake detection using keyframes. Keyframes were extracted from the FaceForensics++ dataset, preserving high-resolution information crucial for robust detection. Each model was trained and tested under identical conditions to ensure fair comparison. The hybrid architecture, combining the local feature extraction capabilities of CNNs with the global contextual modelling power of Transformers, achieved the highest performance across all metrics, including accuracy, precision, recall, F1-score, and AUC. Our findings highlight the superiority of multi-perspective feature learning and reinforce the importance of keyframe utilization in compressed video-based Deepfake detection. This work provides a solid benchmark and foundation for future research on real-time and cross-dataset Deepfake detection frameworks. Key Words: Deepfake Detection, Convolutional Neural Networks (CNN), Transformer Networks, CNN+LSTM, CNN–Transformer Hybrid, Face Forensics++ (FF++), Keyframe Extraction, Deep learning.

Read the paper · More papers on PaperTik