GLAD-ViT: Global–Local Attention Duality in Vision Transformers for Deepfake Detection

Pawan Kumar Pandey, Arun Solanki, Sanjay Kumar Sharma · IETE Technical Review · 2026

The rapid growth of deepfake creation techniques is an important challenge to digital assets and forensic analysis. Recent detection approaches show strong outcomes on visual recognition tasks. However, they exhibit notable limitations in accurately identifying deepfakes. To address these limitations, a novel GLAD-ViT (Global–Local Attention Duality in Vision Transformers) architecture is proposed for enhanced face forgery identification. It incorporates a dual attention mechanism that combines Multi-Head Attention (MHA) with a subsequent Convolutional Block Attention Module (CBAM) within the ViT encoder structure. The global context is modeled by MHA, then CBAM selectively refines local representations through spatial and channel attention mechanisms. The proposed architecture is assessed on several benchmarked deepfake datasets and achieves accuracies of 90.11%, 95.80%, 84.27%, 97.27%, and 99.84% on the DFDC, Celeb-DF (v2), FF++ (F2F), FF++ (FSh), and DF-TIMIT datasets, respectively. Ablation studies confirm the effectiveness of attention-based enhancements across various ViT configurations. Furthermore, the proposed model exhibits robustness against a variety of standard image degradations and significant performance improvement over state-of-the-art methods. This global–local attention duality introduced by GLAD-ViT contributes to advancing deepfake forensics toward more trustworthy media authentication.

Read the paper · More papers on PaperTik