Speed as an Instrument: Exploiting Time-Scale Modification for Adversarial Attacks and Defenses in Speaker Recognition Systems
Umang Patel, Shruti Bhilare, Avik Hati · 2025
Speaker recognition systems (SRS) are important in applications such as critical security systems and virtual assistants. However, deep learning-based SRS models are vulnerable to adversarial attacks that can compromise their reliability, highlighting the need for effective defense strategies. In this paper, we explore time scale modification (TSM) as both a black-box attack and a defense mechanism for SRS models. Specifically, we investigate the effects of the speed-down and speed-up processes on audio signals. Speed-down distorts temporal characteristics, transforming clean samples into adversarial examples (AEs), while speed-up disrupts adversarial patterns, reducing the adversarial attack success rate (AAS R). We evaluate the performance of two SRS models, TDNN and ResNet-34 across different speed manipulations on VoxCeleb and LibriSpeech datasets. Adversar-ial attacks, including FoolHD and PGD, are tested in conjunction with these manipulations. Results show that speed-down is highly effective in generating AEs, significantly lowering model accuracy to 0.08%. Meanwhile, the speed-up process reduces the AASR to as low as 12.85 %, demonstrating its potential as a robust defense method. Our findings offer valuable insights into the vulnerabilities and defense mechanisms of SRS through TSM.