APFT: Adaptive Phoneme Filter Template to Generate Anti-Compression Speech Adversarial Example in Real-Time

Yihuan Huang, Yanzhen Ren, Zongkun Sun, Liming Zhai, Jingmin Wang, Wuyang Liu · IEEE Transactions on Information Forensics and Security · 2025

Automatic Speech Recognition (ASR) systems are widely used for speech censoring. Speech Adversarial Example (AE) offers a novel approach to protect speech privacy by forcing ASR to mistranscribe. However, existing speech AE faces two challenges in real-time voice communication scenarios, such as IP telephone, voice chat, or video conference, it cannot be generated in real-time, and its defensive capability is significantly reduced after the essential audio compression for network transmission. In this paper, we proposeAdaptive Phoneme Filter Template (APFT)to address these issues. The key features of APFT include: 1)Phoneme-level Templatesfor universal AE generation in real-time, 2)Filter, which eliminates redundant signals to improve compression robustness. 3)Adaptive Band Filtering, which limits the attack area from the frequency band without affecting the attack effectiveness and improves speech quality. The comprehensive experimental results show that APFT has four advantages: 1) Real-time Generation, with AE generation time below 1.1ms for 1s speech; 2) Compression Robustness, achieving a WER of 0.64 under AAC and Opus codecs; 3) Transferability, with an average WER of 0.72 across datasets and ASR systems; 4) Stealthiness, achieving a MOS of 4.07 for high-quality speech. In addition, the experiment on Telegram voice calls further proves the practical applicability of APFT. The demo of APFT can be obtained in https://yihuan-qaq.github.io/APFT.github.io/.

Read the paper · More papers on PaperTik