A New Statistical Method of Automatic Lexicon Augmentation

Zheng Y. Zhou · Zhongwen xinxi xuebao · 2001

The out of vocabulary problem is one of the bottlenecks in Chinese Language Modeling.The problem is especially serious for domain specific training data set.This paper presents a new statistical method to extract new words from the training data.This new method is based on association norm estimation,and searches for the word boundaries by right boundary expanding.Combining the new method with another word merging method,we can iteratively optimize the lexicon,segmentation and language model.And very encouraging results are reported in our experiments.

Read the paper · More papers on PaperTik