Generating Targeted Universal Adversarial Perturbation against Automatic Speech Recognition via Phoneme Tailoring

Yujun Zhang, Yanqu Chen, Jiakai Wang, Jin Hu, Renshuai Tao, Xianglong Liu · 2025

There is a growing concern about adversarial attacks against automatic speech recognition (ASR) systems. Although research into targeted universal adversarial examples (AEs) has progressed, current methods are constrained by inefficient exploitation of audio features, demonstrating insufficient attack ability and robustness in the physical world. To solve this problem, we propose a phoneme-tailored attack (PTA) to improve the quality of the generated AEs. Specifically, to improve attack ability, we propose a Diverse Audio Composition Enrichment method, which enhances the utilization of audio features through phoneme-level slicing and recombination. To adapt AEs to complex environments, we propose a Natural Noise Pattern Guidance method to align AEs with natural noise patterns to improve their robustness. Experiments show that our method achieves an average accuracy of more than 72.34% and 98% with and without a norm constraint, and also demonstrates excellent performance in terms of generalization across datasets and resilience to MP3 compression.

Read the paper · More papers on PaperTik