Boosting Deepfake Video Detection Using Non-Local Attention and Sharpness-Aware Minimization
Pratibha Singh, Vinod Pankajakshan · 2025
With the rapid progress and advancement in artificial intelligence, creating synthetic media and manipulating original data has become an incredibly easy task. This poses significant challenges to the authenticity and security of the digital content. Deepfakes are realistic-looking images or videos created using deep learning techniques with the intention to deceive viewers, spread misinformation, influence public opinion, and facilitate financial fraud. While several deepfake countermeasures prove to be effective, they struggle to generalize when it comes to unseen forgeries. In this paper, we present a robust deepfake detection technique that combines ResNet-50 as a feature extraction backbone with a Non-Local Attention Block (NLAB) to enhance spatial dependencies. To improve generalization, we incorporate a Sharpness-Aware Minimization (SAM) technique, which helps the model converge to flat minima, thereby increasing its robustness to small input perturbations and unseen manipulations. In our experiments, the FaceForensics++ and Celeb-df datasets are employed. When trained on the FaceForensics++, our model attains an AUC of 99.11% on the same dataset and generalizes well to Celeb-df with an AUC of 81.39%, surpassing state-of-the-art deepfake detection techniques. Extensive experiments show the effectiveness of the proposed method in improving the performance and generalization of the deepfake detectors.