Analysis of smoothing methods for language models on small Chinese corpora
Ming-Chun Liou, Feng-Long Huang, Ming‐Shing Yu, Yih-Jeng Lin · 2014
Data sparseness has been an inherited issue of statistical language models and smoothing method shave been used to resolve the issue of zero count. 20 Chinese language models from 1M to 20M Chinese words of CGW have been generated on small sizes corpus because of worse situation of zero count issue. Five smoothing methods, such as Good Turing and Advanced Good Turing smoothing, including our 2 proposed methods, are evaluated and analyzed on inside testing and outside testing. It is shown that to alleviate the issue of data sparseness on various sizes of language models. The best one among these methods is our proposed YH-B which performs best in all the various models.