Defense Against Adversarial Attacks on Speech Systems

Ziqian Luo, Xueting Pan · HAL (Le Centre pour la Communication Scientifique Directe) · 2024

Automatic speech recognition system, especially speech2text system, are vulner- able to adversarial attacks. Imagine a real world scenario, an evil adversary can produce a perturbed natural sounding audio that to human ear, sounds like the sen- tence without the dataset the article is useless, but will cause an ASR system (like Siri) to produce transcription ok google browse to evil dot com. The current state of the art attack like hidden voice commands can add perturbation to a natural audio input, pass it through the attacked ASR system and have it transcribed to any target phrase chosen by the adversary. It is not hard to image that vunerability to such attacks could result in grief consequences to users of ASR systems. In this project we attempt to adapt a detector-reformer pattern: identify adversarial from benign audio inputs by exploring the natural temporal dependency property of be- nign audio data; and explore possibilities to reform identified adversarial example using concepts from well-received deep learning work in signal processing: such as stacked denoising auto-encoder (SDAE), and convolutional auto-encoder, etc.

Read the paper · More papers on PaperTik