NOISE-BASED AUDIO-VISUAL FUSION FOR ROBUST SPEECH RECOGNITION
Eric K. Patterson, Sabri Gurbuz, Zekeriya TÜFEKCİ, J.N. Gowdy · 2001
A major goal of current speech recognition research is to improve the robustness of recognition systems used in noisy environments. Recent strides in computing technology have allowed considera-tion of systems that use visual information to augment the deci-sion capability of the recognizer, allowing superior performance in these difficult environments. A crucial area of research in audio-visual speech recognition is how to combine the separate modes of information. Late integration, an approach whereby separate audio-based and video-based decisions are made and then combined “late ” in the process, has emerged as one of the simplest yet most effective techniques. Research has suggested that the fusion method for this technique (and similar methods such as multi-stream HMMs) is af-fected somewhat by the level of interfering audio noise. This paper further defines the relationship between data fusion in the presence of audio noise and demonstrates that optimal data fusion can only be performed if both the noise level and type are considered. 1.