Preventing converted speech spoofing attacks in speaker verification
M. J. Correia, Alberto Abad, Isabel M. Trancoso · 2014
Voice conversion (VC) techniques, which modify a speaker's voice to sound like another's, present a threat to automatic speaker verification (SV) systems. In this paper, we evaluate the vulnerability of a state-of-the-art SV system against a converted speech spoofing attack. To overcome the spoofing attack, we implement state-of-the-art converted speech detectors based on short- and long-term features. We propose a new converted speech detector using a compact feature representation and a discriminative modeling approach. We experiment pairing converted speech detectors based on short- and long-term features to improve converted speech detection. The results indicate that the proposed converted speech detector pair outperforms state-of-the-art ones, achieving a detection accuracy of 97.9% for natural utterances and 98.0% for converted utterances. We include the anti-spoofing mechanism in our SV system as a post-processing module for accepted trials and reevaluate its performance, comparing it with the performance of an ideal system. Our results show that the SV system's performance returns to acceptable values, with less than 1.6% equal error rate (EER) change.