Improved speech recognition using adaptive audio-visual fusion via a stochastic secondary classifier
Simon Lucey, S. Sridharan, Vinod Chandran · 2002
The adaptive fusion of video and audio is one of the fundamental pursuits of audio visual speech recognition (AVSR). In this paper the use of a high dimensional secondary classifier on the word likelihood scores from both the audio and video modalities is investigated for the purposes of adaptive fusion. Results are presented that lie above or equal to the boundary of catastrophic fusion across a number of audio noise levels.