STCA-net: spatio-temporal collaborative attention network for deepfake video detection
Jianping Li, Jing Sun, Yanyi Meng, Kexin Xu · Engineering Research Express · 2025
Abstract Deepfake videos are gradually eroding the trust mechanisms of social platforms due to their high degree of realism.Existing spatio-temporal fusion methods mostly rely on linear superposition strategies and have limited adaptability when facing complex forgery scenarios.To enhance the detection capability for complex forgery signals, this paper proposes the Spatio-Temporal Collaborative Attention Network (STCA-Net), which achieves deep coupling between key region identification and dynamic anomaly capture through a two-stage collaborative attention mechanism of ‘spatial guidance for temporal modeling, temporal feedback for spatial enhancement.’ First, multi-scale spatial feature extraction is constructed,followed by the design of an encoder-decoder collaborative attention mechanism to achieve bidirectional information flow.Finally, through multi-granularity temporal modeling, parallel modeling is performed at the frame-level, segment-level, and video-level with adaptive fusion of the output. Evaluated on four datasets—FaceForensics++, Celeb-DF, DFDC, and DeeperForensics—STCA-Net achieves an average AUROC of 94.2%, reaching 97.6% on FaceForensics++ and 91.4% AUROC on the complex multi-forgery scenario DFDC, representing a 1.3 percentage point improvement over existing best methods. Ablation analysis shows that the two-stage collaborative attention contributes a primary performance improvement of 2.8 percentage points, with multi-granularity modeling further enhancing stability. This method provides an effective solution for complex forgery scenarios and advances both theoretical and practical development of spatio-temporal fusion in deepfake detection.