Modified Heterogeneous Auto-Similarities of Characteristic for Voice Spoofing Attack Detection
Fedila Meriem, Fatiha Mokdad, Ait Sadi Karima, El-Kateb Naima · 2025
Voice spoofing detection is a critical research area due to the increasing reliance on Automatic Speaker Verification (ASV) systems in security-sensitive applications. To address this issue, recent research has focused on developing effective countermeasures to differentiate between genuine and spoofed speech. This study proposes a novel approach to voice spoofing detection using spectrogram image, with an emphasis on enhancing low-level features of the Heterogeneous Auto-Similarities of Characteristic (HASC) descriptor. The main novelty is to integrate a binary code image extracted from the Binary Similarity Image Features (BSIF) descriptor as an additive low-level parameter, to form a new texture descriptor, namely, modified Heterogeneous Auto-Similarities of Characteristic (mHASC). The integration of BSIF allows the model to capture finer local texture details within spectrogram images, addressing the limitations of HASC in detecting fine distinctions between bona fide and spoofed speech. Evaluation in the ASVspoof 2017 v2.0 dataset demonstrates that the proposed mHASC-based countermeasure system significantly outperforms the CQCC-GMM baseline, achieving Equal-Error Rates (EER) of 8.22% and 15.70% on the development and evaluation sets, respectively.