Noisy Speech Recognition Based on Robust End-point Detection and Model Adaptation
Zhipeng Zhang, Sadaoki Furui · 2006
How to detect speech periods in noisy speech and how to cope with the temporal variation of noise characteristics are challenging problems. This paper proposes a new robust noisy speech recognition method based on robust end-point detection and online model adaptation using tree-structured noisy speech HMMs. The basic algorithm consists of; 1) blind speech segmentation; 2) best matching GMM selection; 3) recognizing the speech with the HMM that corresponds to the GMM; 4) end-point detection based on the recognition results; 5) HMM adaptation based on the recognition results; and 6) re-recognition using the adapted HMM. The processes of 1) through 6) are repeated by shifting the blind segmentation window until the end of the sequence of utterances is detected. The proposed method is evaluated by noisy speech collected by a Japanese dialogue system. Experimental results show that the proposed method is effective in recognizing noisy speech under various noise conditions.