On an Empirical Study of Smoothing Techniques for a Tiny Language Model
Fréha Mezzoudj, Mourad Loukam, Abdelkader Benyettou · 2015
The language models (LM) are an important module in many areas of natural language processing, in particular speech recognition and machine translation. In this experimental work, we present the most popular smoothing methods and their effects on statistical language modelling. We compare the behavior of twelve smoothing algorithms that have been developed in speech and natural language processing fields, using a small but novel text corpus of French radio show transcription to construct and improve tiny language models. The perplexity (average word branching factor), which measures the performance of our LM, ranked from 195.9 to 165.4. The best result is obtained by Modified Kneser-Ney algorithm, with the interpolation version. The details of the experimentation are given. We consider the obtained results good and in agreement with the literature.