Integrating Self-Supervised Pre-Training With Adversarial Learning for Synthesized Song Detection
Yankai Wang, Yuxuan Du, Dejun Zhang, Rong Zheng, Jing Deng · 2024
Existing spoofing detection systems often perform poorly when applied to highly realistic synthetic song datasets. To address it, we propose a method that integrates self-supervised pre-training with adversarial learning. Initially, we utilize wav2vec 2.0 to extract audio representations, which are subsequently fed into a back-end classifier. A ResBlock-based network is then employed to capture fine-grained audio features. Additionally, we enhance a RawNet-based model by incorporating a gradient reversal layer and applying adversarial training to improve generalization to unknown algorithms. Finally, the outputs of various models are combined at the score level. Experimental results demonstrate that our approach achieves Equal Error Rates (EER) of 1.57% on the test sets of the Controlled Singing Voice Deepfake Detection (CtrSVDD) track, providing relative reductions of 84.89% compared to the baseline B02. Our system achieves the first place during the CtrSVDD track of SVDD challenge 2024.