Maximum entropy language modeling and the smoothing problem
Sebastián Martín, Hermann Ney, Christoph Hamacher · IEEE Transactions on Speech and Audio Processing · 2000
This paper discusses various aspects of smoothing techniques in maximum entropy language modeling. This topic is typically not addressed in literature. The results can be summarized in four statements: 1) straightforward maximum entropy models with nested features, e.g., tri-, bi-, and uni-grams, result in unsmoothed relative frequencies models, 2) maximum entropy models with nested features and discounted feature counts approximate backing-off smoothed relative frequencies models with Kneser's advanced marginal back-off distribution. This explains some of the reported success of maximum entropy models in the past. 3) We give perplexity results for nested and nonnested features, e.g., trigrams and distance-trigrams, on a 4 million word subset of the Wall Street Journal Corpus. From these results we conclude that the smoothing method has more effect on the perplexity than the method of how to combine the different types of features. 4) We show perplexity results for nonnested features using log-linear interpolation of conventionally smoothed language models, giving evidence that this approach may be a first step to overcome the smoothing problem in the context of maximum entropy.