SASDN: A Generalizable and Minimal-Intervention LLM-Integrated Framework for Continual Adaptation in Spoofed Speech Detection
Utkarsh Venaik, Akash Kushwaha, Nabeel Koya A, Rajiv Ratn Shah · 2025
The rapid rise of generative AI (Gen-AI) technologies has escalated spoofing risks for automatic speaker verification (ASV) systems. Tra- ditional acoustic-only methods often fail when faced with unseen or sophisticated attacks. We propose the Semantic-Aided Spoofing Detection Network (SASDN), which augments RawNet2 with large language model (LLM) support at inference time via a two-step mechanism: (1) a learned embedding space formed during training, and (2) final score fusion weighted by a tunable factor δ. Rather than relying on complex multi-modal training pipelines, SASDN incorporates a novel label impurity strategy that conditions the model to accept uncertain supervision by substituting a fraction of ground truth labels with ''Don't Know.'' At inference, the LLM evaluates transcribed speech against external knowledge to verify authenticity and adjusts the ASV model's prediction accordingly. If the LLM is uncertain, SASDN reverts to the original acoustic classification, ensuring no performance degradation. Evaluations on a VoxCeleb2-derived spoofed dataset demonstrate that SASDN reduces the equal error rate (EER) from 22.71% to 6.5%, offering a continually adaptive defense against emerging threats.