Combining acoustic and visual classifiers for the recognition of spoken sentences

Keren Yu, Xiaoyi Jiang, Horst Bunke · 2002

Acoustic and visual signals carry complementary information and a combination of both information sources therefore possesses the potential of increasing the performance of speech recognition, particularly in noisy environments. In this paper we consider such a combination. Earlier works on the combination of visual and acoustic classifiers for speech recognition typically deal with small vocabularies and use simple combination rules such as majority vote and Borda count. The large number of spoken sentences, however, necessitates a conceptually new approach to classifier combination which explores the syntactic structural of a sentence. In this paper we present such a structure combination strategy and show results for the task of e-mail command recognition.

Read the paper · More papers on PaperTik