Microphone array signal processing for far-talking speech recognition
Jen‐Tzung Chien, Jain-Ray Lai, Po-Yin Lai · 2002
This paper presents a combined microphone array and model adaptation algorithm for distant speech recognition. We aim at resolving the inconvenience of using a head-mounted/hand-holding microphone in a conventional speech recognizer. To improve the distant speech quality, a linear microphone array is applied and acts as a robust acquisition system. We develop a time-domain coherence measure (TDCM) to precisely detect the time delay of speech signals collected by different microphones. The estimated delay is adopted in a delay-and-sum beamformer for speech enhancement. Further, we adapt the speech hidden Markov models to get close to the acoustic condition of enhanced test speech for robust speech recognition. In acquisition and recognition experiments on connected Chinese digits, we find that TDCM can estimate the time delay as precisely as that calculated assuming the speech source direction is known. Increase of speech sampling rate is helpful to determine time delay. Also, the incorporation of the model adaptation scheme can significantly reduce the recognition errors with moderate computation overhead.