Protection Method based on Multiple Sub-Detectors against Audio Adversarial Examples

Keiichi Tamura, Hajime Ito · 2021

Applications with audio speech recognition usually involve personal and authentication information; therefore, security measurement for audio speech recognition is one of the most important issues. Audio adversarial examples are a safety threat in the real world such that automatic speech recognition systems convert an input voice sound to an incorrect text message. Audio adversarial examples deceive the automatic speech recognition systems and they are created by exploiting the vulnerabilities of the deep-learning-based automatic speech recognition systems with the speech-to-text transcription neural network technique. To protect applications with automatic speech recognition, a new method based on multiple sub-detectors against audio adversarial examples is proposed in this study. The proposed protection method has a detector that is composed of three sub-detectors: dynamic-sampling-based, denoising-based, and temporal-dependency-based sub-detectors. In the experiments, we created 1,000 audio adversarial examples for evaluation. These audio adversarial examples are created on Mozilla-implemented Deep Speech. Deep Speech is a deep-learning-based end-to-end ASR system and it is the main attack target of this study. The experimental results of detecting audio adversarial examples show that the protection performance of the new method is higher than those of conventional methods.

Read the paper · More papers on PaperTik