LiveGuard: Voice Liveness Detection via Wavelet Scattering Transform and Mel Spectrogram Scaling
Liqun Shan, Xingli Zhang, Md Imran Hossen, Xiali Hei · 2025
Voice-controlled interfaces are essential in modern smart devices, but they remain vulnerable to replay attacks that compromise voice authentication systems. Existing voice liveness detection methods often struggle to distinguish human speech from replayed audio. This paper introduces a novel approach, LiveGuard, utilizing wavelet scattering transform (WST) and Mel spectrogram scaling with a lightweight ResNet architecture to enhance voice liveness detection. WST captures robust hierarchical features, while Mel spectrogram scaling extracts fine-grained acoustic details, which the lightweight ResNet efficiently processes to identify live voice. Experimental results demonstrate accuracy improvements of 6% with WST and Mel spectrogram scaling, achieving a top accuracy of 97.17% on POCO dataset. Meanwhile, LiveGuard demonstrates superior performance on ASVspoof2019 and ASVspoof2021 benchmarks. It achieves the lowest equal error rate (EER) of 0.13%, and a min t-DCF of 0.00126 on ASVspoof2019, and an EER of 0.42% on ASVspoof2021, surpassing state-of-the-art methods.