A Lightweight and Efficient Model for Audio Anti-Spoofing

Qiaowei Ma, Jinghui Zhong, Yitao Yang, Weiheng Liu, Ying Gao, Wing W. Y. Ng · 2023

With the rapid development of speech conversion and speech synthesis algorithms, automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. In recent years, researchers had proposed anti-spoofing systems based on hand-crafted features. However, using hand-crafted features rather than raw waveform will lose implicit information for audio anti-spoofing. Inspired by the promising performance of ConvNeXt in classification tasks, we reference the network architecture design of ConvNeXt and propose a Lightweight and Efficient Model for Audio Anti-Spoofing (LEMAAS). With no preceding feature extraction process, we employ raw waveforms as direct inputs to our proposed model. By integrating with the channel attention module and using the focal loss function, the proposed model can focus on the most informative features representation of speech and the difficult samples that are hard to classify. Experimental results show that our proposed system could achieve an equal error rate of 0.64% and min-tDCF of 0.0187 for the ASVspoof 2019 LA evaluation dataset, which outperforms the state-of-the-art systems. Moreover, even when trained only on the ASVspoof 2019 LA dataset, the model still achieved equal error rates of 0.86% and 1.18% on the ASVspoof 2015 development dataset and evaluation dataset, respectively. This demonstrates that our model has achieved promising generalization performance during cross-dataset testing.

Read the paper · More papers on PaperTik