n-gram estimates in probabilistic models for Pinyin to Hanzi transcription

Amelia Fong Lochovsky, Hon-kit Cheung · 2002

We consider the problem of sparse data in probabilistic modeling of the Chinese language. To date, n-gram models outperform models that try to capture linguistical structures. Various techniques for estimating n-gram statistics for the English language have been proposed and compared. It is known that how various techniques actually perform depends on the problem domain in which the probabilistic model is applied. We apply different smoothing techniques in the estimates of bigram statistics in a word based bigram model for Pinyin to Hanzi transcription. Comparative results are reported and show improved accuracy over the MLE method. We have also experimented with hybrid approaches (using bigrams as well as monograms) to achieve superior results.

Read the paper · More papers on PaperTik