Transformer-Based Ensemble Approach for Robust Deepfake Detection in Low-Resolution Short Videos
Chun Yun Chiang, Hui Na Chua, Muhammed Basheer Jasser, Richard T.K. Wong, Bayan Issa · 2025
Deepfake technology—synthetic media that replaces a person's likeness in existing images or videos—has raised significant concerns in politics, security, and the media. Recent advances such as ChatGPT 4.0 have further magnified the potential misuse of these synthetic artifacts, making it increasingly challenging to detect maliciously manipulated short videos. This paper proposes a deepfake detection framework to address challenges in low-resolution, compressed, and short video scenarios. The methodology leverages Haar Cascade for face detection, a Vision Transformer (ViT) for feature extraction, and an ensemble classifier (Support Vector Machines, Random Forest, Logistic Regression) for final data modelling. Our result findings underline how high compression, video distortions, and limited training data affect detection accuracy. Experiments on Celeb-DF (v2) and DeeperForensics-1.0 datasets reveal that our proposed method balances computational efficiency and robust detection, outperforming cutting-edge Convolutional Neural Network-based models (e.g., MesoNet) on selected metrics and data partitions. Subsequently, our findings revealed that the proposed ensemble model achieves high accuracy of 74% in detecting the deepfake content on the Celeb-DF (Version 2) and DeeperForensics-1.0 datasets. By focusing on spatial anomalies and practical constraints such as short video duration, this work contributes to a valuable understanding of these constraints and the scalability of the existing approaches.