Unsupervised Word Segmentation Without Dictionary

Jason S. Chang, Tracy Lin · 2003

This prototype system demonstrates a novel method of word segmentation based on corpus statistics. Since the central technique we used is unsupervised training based on a large corpus, we refer to this approach as unsupervised word segmentation. The unsupervised approach is general in scope and can be applied to both Mandarin Chinese and Taiwanese. In this prototype, we illustrate its use in word segmentation of Taiwanese Bible written in Hanzi and Romanized characters. Basically, it involves: Computing mutual information, MI, between Hanzi and Romanized characters A and B. If A and B have a relatively high MI, we lean toward treating AB as a word. Using a greedy method to form words of 2 to 4 characters in the input sentences. Building an N-gram model from the results of first-round word segmentation Segmenting words based on the N-gram model Iterating between the above two steps: building N-gram and word segmentation Computing mutual information. Using mutual information is motivated by the observation of previous work by Hank and Church (1990) and Sproat and Shih (1990). If A and B have a relatively high MI that is over a certain threshold, we prefer to identify AB as a word over those

Read the paper · More papers on PaperTik