Good-Turing estimation from word lattices for unsupervised language model adaptation
Michael Riley, Brian Roark, Richard Sproat · 2004
We present a comparison of using the weighted word lattice output of a recognizer versus its one-best transcription for unsupervised language model adaptation. We begin with a general analysis of how to smooth word probabilities when the sample is hidden, as is the case with recognizer lattices. For each smoothing technique for the known sample case, we show there is a natural generalization to the hidden case. In particular, we use this generalization with the well-known Good-Turing estimate on word lattices, and show results using Monte Carlo methods for building Katz backoff models. In our realistic adaptation task, with mismatched acoustic and language models, we find that Katz backoff models trained on word lattice samples provide a small, consistent benefit over those trained on one-best output, most notably when there is a limited amount of adaptation data (less than 100 hours). Thus, while the recognizer one-best transcription can provide an effective approximation for the purpose of language model adaptation under certain circumstances, the word lattice provides information that can be exploited for more robust language modeling.