A Unified Framework for Detecting Diverse Voice Spoofing Attacks

Parth Rajeshbhai Bhanderi · 2025

Voice spoofing attacks, including logical access (LA) and physical access (PA), pose significant challenges to automatic speaker verification (ASV) systems. Existing countermeasures often target a single attack type, leading to increased computational complexity and limited applicability in real-world scenarios where the nature of the spoof attack is unknown. This paper proposes UNIGUARD, a unified framework designed to detect multiple spoofing attack types. Leveraging the ESResNeXt AudioCLIP encoder, UNIGUARD incorporates band-specific processing across frequency bands and a 2D attention mechanism with depthwise separable convolutions to efficiently capture diverse spoofing artifacts, thereby enhancing its ability to detect various attack patterns. Evaluated on the ASVspoof 2019 dataset, UNIGUARD obtains 7.72% EER for LA, 6.10% for PA, and 5.33% overall, with a 33.4% improvement in minimum tDCF over existing unified approaches.

Read the paper · More papers on PaperTik