Frequency-Domain Masking and Spatial Interaction for Generalizable Deepfake Detection
Xinyu Luo, Yu Wang · Electronics · 2025
Over the past few years, the rapid development of deepfake technology based on generative models has posed a significant threat to the field of information security. Despite the notable progress in deepfake-detection methods based on the spatial domain, the detection capability of the models drops sharply when dealing with low-quality images. Moreover, the effectiveness of detection relies on the realism of the forged images and the specific traces inherent to particular forgery techniques, which often weakens the models’ generalization ability. To address this issue, we propose the Frequency-Domain Masking and Spatial Interaction (FMSI) model. The FMSI model innovatively introduces masked image modeling in frequency-domain processing. This prevents the model from focusing too much on specific frequency-domain features and enhances its generalization ability. We design a high-frequency information convolution module for spatial and channel dimensions to help the model capture subtle forgery traces more effectively. Also, we creatively design a dual stream architecture for frequency-domain and spatial-domain information interaction and overcome single-domain detection limitations. Our model is tested on three public benchmark datasets (FaceForensics++, Celeb-DF, and WildDeepfake) through intra-domain and cross-domain experiments. The detection and generalization capabilities of the model are evaluated using the AUC and EER metrics. The experimental results demonstrate that our model not only possesses high detection capability but also exhibits excellent generalization ability.