Integrating Self-Supervised Pre-Training With Adversarial Learning for Synthesized Song Detection

Yankai Wang, Yuxuan Du, Dejun Zhang, Rong Zheng, Jing Deng · 2024

Existing spoofing detection systems often perform poorly when applied to highly realistic synthetic song datasets. To address it, we propose a method that integrates self-supervised pre-training with adversarial learning. Initially, we utilize wav2vec 2.0 to extract audio representations, which are subsequently fed into a back-end classifier. A ResBlock-based network is then employed to capture fine-grained audio features. Additionally, we enhance a RawNet-based model by incorporating a gradient reversal layer and applying adversarial training to improve generalization to unknown algorithms. Finally, the outputs of various models are combined at the score level. Experimental results demonstrate that our approach achieves Equal Error Rates (EER) of 1.57% on the test sets of the Controlled Singing Voice Deepfake Detection (CtrSVDD) track, providing relative reductions of 84.89% compared to the baseline B02. Our system achieves the first place during the CtrSVDD track of SVDD challenge 2024.

Read the paper · More papers on PaperTik