COMBINING NOISE COMPENSATION WITH VISUAL INFORMATION IN SPEECH RECOGNITION
Stephen Cox, Iain A. Matthews, Jenny Bangham · 1997
The addition of visual information derived from the speaker’s lip movements to a speech recogniser (speechread-ing) can significantly enhance the performance of the recog-niser when it is operating under adverse signal-to-noise ra-tios. However, processing of video signals imposes a large computational demand on the system and there is little point in using speechreading techniques if similar performance gains can be obtained using techniques which operate on only the audio signal and which are less computationally ex-pensive. In this paper, we show that combining visual infor-mation with an audio noise compensation technique (spectral subtraction) leads to a performance significantly higher than that obtained using speechreading only or noise compensa-tion only. The optimum method for speech recognition in the presence of noise is to use speech models that are matched to the input speech, and we show that the addition of visual in-formation also gives a performance gain when matched mod-els are used. We also describe a method of “late ” integration which uses a measure of confidence derived from informa-tion output by the audio recogniser to achieve a performance which is close to optimum. 1.