Speaker/speech recognition using microphone arrays and neural networks
Qiguang Lin, ChiWei Che, Ea-Ee Jan, James L. Flanagan · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1994
Hands-free operation of speech processing equipment is sometimes desired so that the user is unencumbered by hand-held or body-worn microphones. This paper explores the use of array microphones and neural networks (MANN) for robust speech/speaker recognition in a reverberant and noisy environment. Microphone arrays provide high-quality, hands-free sound capture at distances, and neural network processors compensate for environmental interference by transforming speech features of the array input to those of close-talking microphone input. The MANN system is evaluated using both computer-simulated degraded speech and real- room collected speech. It is found that the MANN system is capable of elevating recognition accuracies under adverse conditions, such as room reverberation, noise interference, and mismatch between the training and testing conditions, to levels comparable to those obtained with close-talking microphone input under a matched training and testing condition.