Ambiguity Resolution in Chinese Word Segmentation

Maosong Sun, Benjamin K. Tsou · Waseda University Repository (Waseda University) · 1995

A new method for Chinese word segmentation named Conditional FB.. ►\-1.F. ffi flf.91±. At least two possible segmentations which overlap in position exist for the sequence IP, and ffi if we simply match it with a Chinese dictionary. • Type II — Categorial Ambiguity (CA) (2)a. VT Eljf kittf4 b. ±. Note the sequence --T-1 in (2): the constituents should be combined as a single word in (a), but they should be separated in (b) because of the productivity of expressions such as 1j/17, Basic methods for dealing with segmentation ambiguities so far can be either rule-based [2,3] or statistics-based [4,5]. The former employs the conventional maximal matching strategy, either forward or backward (referred here as RIM and BMM respectively), or both , as a detector of ambiguities, and applies relevant rules from the rule base to solve them. The problems in this approach are (i) the constructions of 0As in texts to be processed are nearly unpredictable, resulting in unwieldy complexity in rule base establislunent and maintenance, and (ii) it always fails in finding CAs. The latter gives segmentation possibilities exhaustively by dictionary lookup as a * This research is supported in part by the Youth Science Foundation of Tsinghua University, Beijing, and by the Language Information Sciences Research Centre, City University of Hong Kong

Read the paper · More papers on PaperTik