A Speech Spoofing Detection Method Based on Feature Fusion and Residual Attention

Haiyan Lan, L. Y. Dong, Yuhua Wang, Yike Li, Dongxu Jiang, Zhiying Han · 2024

The emergence of various types of deception attacks presents a substantial risk to Automatic Speaker Verification (ASV) systems, and an increasing number of experts and scholars are focusing on the research of voice deception detection. This paper introduces an innovative approach for detecting voice deception that addresses the accuracy and generalization limitations of traditional methods. The introduced method is based on spectral-temporal feature fusion, which complements the acoustic features captured by different channels to obtain more comprehensive speaker information. It then utilizes multi-channel fused spectrogram features, applying a residual attention learning mechanism to assign varying levels of attention to the spectrogram features from different channels. The method further leverages an improved ResNet network to learn more refined acoustic information for classification. The method attained an EER of 3.60 and a t-DCF score of 0.0931 on the LA logical access corpus of ASVspoof 2019, which represent improvements of 55.5% and 56.0%, respectively, over the best performing baseline system.

Read the paper · More papers on PaperTik