Language model adaptation for conversational speech recognition using automatically tagged pseudo-morphological classes

C. Crespo, Daniel Tapias, G. Escalada, J. Alvarez · 2002

Statistical language models provide a powerful tool for modelling natural spoken language. Nevertheless a large set of training sentences is required to estimate reliably the model parameters. The authors present a method for estimating n-gram probabilities from sparse data. The proposed language modeling strategy allows one to adapt a generic language model (LM) to a new semantic domain with just a few hundred sentences. This reduced set of sentences is automatically tagged with eighty different pseudo-morphological labels, and then a word-bigram LM is derived from them. Finally, this target domain word-bigram LM is interpolated with a generic back-off word-bigram LM, which was estimated using a large text database. This strategy reduces by 27% the word error rate of the SPATIS (SPanish ATIS) task.

Read the paper · More papers on PaperTik