A robust unified spoofing audio detection scheme
Hao Meng, Wei Ou, Ju Huang, Haozhe Liang, Wenbao Han, Qionglu Zhang · Computers & Electrical Engineering · 2024
The rapid development of generative artificial intelligence and deepfakes techniques makes it increasingly tough to identify different kinds of spoofing audio, which on one hand weakens the safety of Automatic Speaker Verification (ASV) systems, and on the other hand, negatively affects national security and societal stability. In the face of the emergence of constantly evolving forgery techniques, we need a detection model with strong generalization and high robustness to deal with them. In this work, we propose a Robust and Unified Spoofing Audio Detection (RUSAD) scheme, which is capable of dealing with multiple attacks such as logic attacks, replay attacks and adversarial attacks. We propose an innovative scheme from the perspectives of upstream feature extraction and downstream spoofing classification: the self-supervised upstream learning model based on the Conformer network extracts speech representations, establishes probabilistic spectrum-augmentation events to improve model robustness against adversarial attacks, and formulate multiple decoding tasks for classification and regression; downstream classification of audio spoofing based on inference of the SE-ResNeXt network, supplemented by self-attention pooling and the OC-Softmax angular loss function in order to improve classification performance. We perform detailed experimental evaluations on both the ASVspoof2021 and ADD2023 datasets. The results show that the scheme has improved performance, generalization, and robustness in comparison to the baseline system as well as to most forensic countermeasures, which are capable of achieving the goal of dealing with multiple attacks.