Language-model optimization by mapping of corpora

Dietrich Klakow · 2002

It is questionable whether words are really the best basic units for the estimation of stochastic language models-grouping frequent word sequences to phrases can improve language models. More generally, we have investigated various coding schemes for a corpus. In this paper, it is applied to optimize the perplexity of n-gram language models. In tests on two large corpora (WSJ and BNA) the bigram perplexity was reduced by up to 29%. Furthermore, this approach allows to tackle the problem of an open vocabulary with no unknown word.

Read the paper · More papers on PaperTik