The Impact of Data Augmentation on Deepfake Detection: A Comparative Study of CNN and Transformer Models

Noor Hayder Abdul Ameer, Mustafa M. Al-Taee, Abadal-Salam T. Hussain, Ahmed Hussian, Ali Haider · 2025

As AI becomes increasingly prevalent in our daily lives, the rise of deepfake technology presents a significant challenge for digital forensics and security, necessitating the development of robust detection models to help distinguish between real and fake media. In this work, We study the effect of data augmentation on deepfake classification performance using CNN-based and Transformer based networks. XceptionNet, ResNet-50, EfficientNet-B3, ViT-B/16, and ConvNeXt-Small are evaluated on the Deepfake Detection Challenge (DFDC) dataset and augmentation operations consist of rotating, color jitter, Gaussian noise and random erasing. The results also indicate that CNN-based models exceed the performance of Transformer-based architectures, as XceptionNet achieves an accuracy score of 94% and an F1-score of 0.94, while ViT-B/16 performs poorly, achieving an accuracy of 82%. These show that augmentation helps improve robustness for CNN models, while Transformer models show a less effective performance in this setup. The paper offered guidance on optimizing data preprocessing approaches for deepfake identification and proposed further investigation of hybrid CNN-Transformer architectures that could empower models to be more effective in detecting fake images.

Read the paper · More papers on PaperTik