An Attention Fusion-Based Multimodal Autoencoder Deepfake Detection Framework
Tong Hu, Wei Liu, Yaqi Fan, Jie Yang, Yuanyuan Qiao · 2025
Deepfake detection is a critical task to mitigate the risks posed by AI-generated fake media content across societal, political, and security domains. As the quality of generated content improves, traditional single-modal detection methods increasingly struggle to address the complexity and diversity of deepfake content. To overcome the limitations of single-modal analysis and tackle challenges such as large-scale datasets and high annotation costs, we propose an Attention Fusion-based Multimodal AutoEncoder (AFM-AE) Framework. This paper designs a dual-branch autoencoder detection network that integrates both video and audio modalities, innovatively incorporating multi-level attention mechanisms to enable efficient feature fusion. Additionally, we propose a tailored loss function and training enhancement strategies to optimize initially weakly-correlated multimodal features. Experimental results demonstrate that the method significantly improves detection accuracy within an unsupervised learning framework, achieving an AUC of 80.13% on the large public DFDC dataset, comparable to supervised methods on the same dataset. This approach provides a solution with low data-labeling dependency, high efficiency, and strong practical applicability.