A study of mutual front-end processing method based on statistical model for noise robust speech recognition

Masakiyo Fujimoto, Kentaro Ishizuka, Tomohiro Nakatani · 2009

This paper addresses robust front-end processing for automatic speech recognition (ASR) in noise. Accurate recognition of corrupted speech requires noise robust front-end processing, e.g., voice activity detection (VAD) and noise suppression (NS). Typically, VAD and NS are combined as one-way processing, and are developed independently. However, VAD and NS should not be assumed to be independent techniques, because sharing each others’ information is important for the improvement of front-end processing. Thus, we investigate the mutual front-end processing by integrating VAD and NS, which can beneficially share each others’ information. In an evaluation of a concatenated speech corpus, CENSREC-1-C database, the proposed method improves the performance of both VAD and ASR compared with the conventional method. Index Terms: voice activity detection, noise suppression, mutual front-end processing, speech recognition

Read the paper · More papers on PaperTik