Audio Steganography Based Backdoor Attack for Speech Recognition Software

Mengyuan Zhang, Shunhui Ji, Hanbo Cai, Hai Ying Dong, Pengcheng Zhang, Yunhe Li · 2024

With the growing prevalence of deep learning in the speech area, speech recognition, voice control, and related applications have become integral parts of people's lives. However, the rise of malicious third-party platforms has introduced significant security concerns, particularly through backdoor attacks. These attacks implant triggers that manipulate speech recognition models to produce specific labels, thereby compromising the system's integrity. Studying speech backdoor attacks is crucial for evaluating the security of speech recognition software, and iden-tifying and addressing potential vulnerabilities. Existing methods for speech backdoor attacks usually employ fixed perturbations as triggers. However, these perturbations may be discernible to the human ear, making them easily detectable. To address this issue, we propose a frequency domain-embedded backdoor attack method based on echo hiding. Echo hiding is a steganography technique based on audio. This method embeds hidden information into the frequency spectrum of the echo signal, leveraging the masking property of the human auditory system. It is difficult to arouse suspicion or detect the presence of hidden information since echo is perceived as a natural phenomenon in auditory perception. Furthermore, it does not cause a significant decrease in audio quality. Experimental results show the effectiveness of our method in different settings.

Read the paper · More papers on PaperTik