Using likelihood L-statistics to measure confidence in audio-visual speech recognition
A. Ghosh, Ashish Kumar Verma, Achintya Kumar Sarkar · 2002
This paper describes previous work on decision fusion in audio-visual speech recognition. A novel approach is proposed to combine audio and video channel information in audio-visual speech recognition scenario. We have considered frame-level phonetic classification problem using two single-stream Gaussian mixture models. Audio and video streams are adaptively weighted using a cumulative mean of the sample confidence values over past frames in addition to the present sample confidence value. The confidence values for audio and video decisions are computed using an L-statistics (linear combination of order-statistics) of log-likelihoods against phone models. It is shown through various experiments, on a database of about 15000 sentences from large vocabulary continuous speech, that the proposed approach results in better classification accuracy as compared to other approaches.