A voice spoofing Detection Model based on Dilated Residual Attentional Feature-fusion Net with Enhanced Feature Extraction
Wenxuan Xu, Xiao Zou, Dong Hu, Shengyou Qian · 2024
The rapid advancement of science and technology in recent years has elevated the threat posed by voice spoofing techniques, including Text To Speech, voice conversion, imitation, and replay to Automatic Speaker Verification (ASV) systems. Detecting the authenticity of speech input to ASV systems and mitigating the risk of voice spoofing attacks have become critical issues in the field of speech research. This paper introduces various types of voice spoofing and the framework for voice spoofing detection. Additionally, it proposes a Dilated Residual Attentional Feature-Fusion Net (DRAFN) voice spoofing detection model to address the challenge of insufficient speech feature extraction. Experiments were conducted using the LA dataset from the ASVspoof 2021 competition, and the results demonstrated that the EER and t-DCF of the DRAFN model decreased by 5.5% and 5.8% compared to the baseline models of voice spoofing detection at least. These findings underscore the effectiveness of the DRAFN model in enhancing the robustness of ASV systems against voice spoofing attacks.