Exploring Deepfake Detection: A Comparative Analysis

Rancy Chepchirchir, Julius Sechang Mboli · 2025

Deepfakes-synthetic media generated using deep learning techniques such as generative adversarial networks (GANs), autoencoders, and diffusion modelspose threats to digital trust, especially in political, financial, and personal contexts. Despite advancements in deepfake detection, existing models often fail to generalize across different manipulation techniques and datasets, limiting their real-world applicability. This study aims to evaluate and compare the effectiveness of three deep learning architectures-Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Vision Transformers (ViTs)-in detecting deepfakes across three benchmark datasets: FaceForensics++, Celeb-DF, and the DeepFake Detection Challenge (DFDC) dataset. We implemented individual and hybrid models, extracted spatial and temporal features, and assessed model performance using accuracy, precision, recall, F1-score, and ROC-AUC, along with cross-dataset evaluation to test generalizability. We found that ViTs outperform CNNs and RNNs, achieving up to 84.22% accuracy on both DFDC and Celeb-DF datasets. In contrast, CNNs and RNNs exhibited strong performance on seen data but suffered from performance drops when applied to unseen manipulations. These findings highlight the superior generalization ability of ViTs and emphasize the importance of hybrid and Transformer-based architectures for developing robust, real-world deepfake detection systems.

Read the paper · More papers on PaperTik