From Voices to Beats: Enhancing Music Deepfake Detection by Identifying Forgeries in Background

Zi Wei, Dengpan Ye, Jiacheng Deng, Yuhan Lin · 2025

Music deepfake detection is aimed at identifying whether songs are generated by AI. Current methods usually separate vocals from background music for detection, but this could leave residual forgery information in the background. Our study demonstrates for the first time that incorporating background forgery information with vocals can improve detection accuracy. Furthermore, our findings show that using background features can reduce EER by an average of about 2% on existing frameworks. Based on this observation, we propose a novel Hybrid Frontend that captures generalized features from both vocal and background music. The Hybrid Frontend comprises two branches: vocal and background music part. Specifically, the vocal part uses the Sinconv encoder as a deeply embedded feature extractor. The latter captures background variation by fine-tuning the pre-trained model with adapters. Experimental results demonstrate that our method outperforms the vocal-only detection on WildSVDD dataset, achieving an EER of 8.53%, which is 1.3% lower.

Read the paper · More papers on PaperTik