Adapted language modeling for recognition of retelling story in language learning
Meng Chen, Yang Song, Lan Wang · 2012
N-gram language modeling typically requires large quantities of in-domain training data, i.e., data that matches the task in both topic and style. For the task of retelling stories, obtaining large volumes of speech transcriptions is often unrealistic. In this paper, we propose a novel method of language modeling using mixture models with very limited text datain the task of retelling stories. We modeled topic-specific, spoken-style, and document-style language models separately and interpolated them. We also interpolated the class-based language model with the N-gram models. Experimental results show that up to 61.6% reduction of perplexity and 20.7% reduction of word error rate (WER) have been obtained by our best performing model.