Backoff inspired features for maximum entropy language models
Fadi Biadsy, Keith Hall, Pedro J. Moreno, Brian Roark · 2014
Maximum Entropy (MaxEnt) language models [1, 2] are linear models that are typically regularized via well-known L1 or L2 terms in the likelihood objective, hence avoiding the need for the kinds of backoff or mixture weights used in smoothed n-gram language models using Katz backoff [3] and similar tech-niques. Even though backoff cost is not required to regularize the model, we investigate the use of backoff features in Max-Ent models, as well as some backoff-inspired variants. These features are shown to improve model quality substantially, as shown in perplexity and word-error rate reductions, even in very large scale training scenarios of tens or hundreds of billions of words and hundreds of millions of features. Index Terms: maximum entropy modeling, language model-ing, n-gram models, linear models