CCFormer: A cascaded transformer framework for precise temporal audio-visual deepfake localization
Peifan Li, Jinluan Ren · Alexandria Engineering Journal · 2025
Audio-visual deepfake detection presents significant computational challenges in achieving precise temporal boundary localization beyond traditional binary classification approaches. This study presents CCFormer, a cascaded optimization framework that integrates ConvNeXt-V2 visual forgery detection with CrossFormer cross-modal localization for precise temporal forgery localization. The framework employs a two-stage strategy where ConvNeXt-V2 performs efficient suspicious segment screening through multi-scale spatiotemporal feature extraction, while CrossFormer achieves frame-level precision through multi-head cross-modal attention mechanisms for optimal audio-visual feature alignment. Experiments on the LAV-DF dataset demonstrate that CCFormer achieving 96.30 % [email protected] and 84.96 % [email protected] The framework achieves inference time of 23.4 ms per video, representing 58.1 % improvement over conventional end-to-end architectures. Ablation studies reveal that the CrossFormer module increases detection performance in high-precision IoU intervals by 153.4 % compared to the baseline methods. The optimization framework successfully transforms coarse-grained binary classification into precise temporal boundary localization,