UltraAdv: An Ultrasonic Adversarial Attack on Closed-Box Speech Recognition Systems
Guoming Zhang, Xiaohui Ma, Huiting Zhang, Riccardo Spolaor, Yanni Yang, Xiaoyu Ji, Xiuzhen Cheng, Pengfei Hu · IEEE Transactions on Mobile Computing · 2025
Attacks on speech recognition systems often use adversarial or inaudible commands. However, a challenge is that adversarial perturbations typically fall within the audible frequency range, making it difficult to achieve inaudibility. Additionally, the non-linear effects of loudspeakers often cause inaudible commands to become audible at higher power levels. Therefore, minimizing the power requirements of the attack is essential to maintain inaudibility. Another significant obstacle is the conversion of variable-length commands, especially longer ones, into shorter target commands. In this paper, we present UltraAdv, a method for generating long-range adversarial perturbations capable of compromising commands of arbitrary length in closed-box setting. By combining the ultrasonic signal with the normal one, rather than negating it as in DolphinAttack, we significantly improve the energy efficiency, thus enhancing its attack distance. We also propose a dynamically adjustable suppression-interference method based on automatic gain control to address the challenge of mismatched durations between long commands and target commands (length-independent). Experiments demonstrate that using a single perturbation, we achieve impressive success rates of 98.84% and 96.62% and 98.32% across a diverse set of 12,260 speeches on DeepSpeech, iFlytek, and Whisper. The attack range reaches up to 15 m, surpassing DolphinAttack's 5 m range at equivalent power.