Segmenting Chinese Based on Probabilistic Model

Yafei Zhang · Acta Simulata Systematica Sinica · 2002

Word Segmentation is a basic task of Chinese Information Processing. In this paper we present a simple probabilistic model of Chinese text based on the occurrence probability of the words, which can be seen as a zero-th order hidden Markov Model (HMM). Then we investigate how to discover by EM algorithm the words and their probabilities from a corpus of unsegmented text without using a dictionary. The last part presents a simulation system of processing Chinese text.

Read the paper · More papers on PaperTik