Anti-Spoofing Detection on Zero-Shot Synthetic Audios With Self-Supervised Representation Learning
Ben Niu, Shiqi Zhou, Jiayu Li, Jin Ping Cao, Yuqiao Hou · 2024
Synthetic speech spoofing attacks maliciously use machine-synthesized speech to impersonate individuals and bypass speaker verification systems, posing security risks to people’s privacy. In this paper, we propose an enhanced anti-spoofing detection module, which utilizes self-supervised representations to differentiate between synthetic and bona fide speech, especially on the advanced zero-shot synthetic speech. Experimental results demonstrate the effectiveness of our method, achieving a 65% reduction in the equal error rate metric for detecting zero-shot spoofing, and a 79% enhancement in generalizing against unknown attacks compared to baseline systems. Extensive experiments also validate the efficacy of applying pre-trained models in zero-shot spoofing detection, and demonstrate the value of using state-of-the-art zero-shot synthetic speech data to enhance anti-spoofing datasets.