EVA-ASCA: Enhancing Voice Anti-Spoofing through Attention-based Similarity Weights and Contrastive Negative Attractors
Nghi Tran, Bima Prihasto, Phuong Thi Le, Thao Quoc Tran, Chun-Shien Lu, Jia‐Ching Wang · 2024
Voice spoofing attacks pose an escalating security concern within the contemporary digital landscape. Attackers employ techniques such as voice conversion (VC) and text-to-speech (TTS) to generate a synthetic speech that replicates the victim’s voice, to illicitly access sensitive data. Detection of these attacks hinges on identifying anomalies in audio transmission resulting from these deceptive activities. Anomalies arise from encoding and transmission conditions that are not commonly encountered, particularly in situations such as local authentication or telephony. To address this issue, our study presents a strategy featuring pivotal enhancement: Attention-based Similarity Weights and Contrastive Negative Attractors. This technique clusters authentic speeches around multiple speaker attractors together in a high-dimensional embedding space, effectively thwarting spoofing attacks across all attractors. Experimental results substantiate the superiority of our system, yielding a substantial 1.09% improvement in the equal error rate (EER) when compared to existing solutions on the ASVspoof 2019 LA evaluation dataset.