Automatic transcription of lecture speech using topic-independent language modeling

Kazuomi Kato, Hiroaki Nanjo, Tatsuya Kawahara · 2000

We approach lecture speech recognition with a topicindependent language model and its adaptation. As lecture speech has its characteristic style that is different from newspapers and conversations, dedicated language modeling is needed. The problem is that, although lectures have many keywords specific to the topic and fields, available corpus of each domain is limited in size. Thus, we introduce topic-independent modeling with a vocabulary selection mechanism based on a mutual information criterion. It realizes better coverage and accuracy with small complexity than the conventional word frequency-based method. This baseline model is adapted to specific lectures using preprint texts. We have tried automatic transcription of oral presentations and achieved a word error rate of 23.6% on the average.

Read the paper · More papers on PaperTik