Modified Kneser-Ney Smoothing of n-gram Models

Frankie James · 2000

This report examines a series of tests that were performed on variations of the modified Kneser–Ney smoothing model outlined in a study by Chen and Goodman. [2] We explore several different ways of choosing and setting the discounting parameters, as well as the exclusion of singleton contexts at various levels of the model. Statistical language modeling can be used effectively to provide a baseline for recognition accuracy when studying other forms of speech and language recognition. Perplexities computed using smoothed n-gram models can later be compared to language models based on grammars. In this paper, we look at perplexities calculated on ATIS travel data using the statistical language model known as modified Kneser–Ney. We explore four variations of the basic algorithm outlined in Chen and Goodman [2], and select one that appears to perform significantly better on our test data. We plan to use this model as the baseline for our analysis of future grammar models.

Read the paper · More papers on PaperTik