Chinese Word Auto-Confirmation Agent

Jia‐Lin Tsai, Cheng-Lung Sung, Wen−Lian Hsu · 2003

In various Asian languages, including Chinese, there is no space between words in texts. Thus, most Chinese NLP systems must perform word-segmentation (sentence tokenization). However, successful word-segmentation depends on having a sufficiently large lexicon. On the average, about 3 % of the words in text are not contained in a lexicon. Therefore, unknown word identification becomes a bottleneck for Chinese NLP systems. In this paper, we present a Chinese word auto-confirmation (CWAC) agent. CWAC agent uses a hybrid approach that takes advantage of statistical and linguistic approaches. The task of a CWAC agent is to auto-confirm whether an n-gram input (n ≥ 2) is a Chinese word. We design our CWAC agent to satisfy two criteria: (1) a greater than 98 % precision rate and a greater than 75 % recall rate and (2) domain-independent performance (F-measure). These criteria assure our CWAC agents can work automatically without human intervention. Furthermore, by combining several CWAC agents designed based on different principles, we can construct a multi-CWAC agent through a building-block approach. Three experiments are conducted in this study. The results demonstrate that, for n-gram frequency ≥ 4 in large corpus, our CWAC agent can satisfy the two criteria and achieve 97.82 % precision, 77.11 % recall, and 86.24 % domain-independent F-measure. No existing systems can achieve such a high precision and domain-independent F-measure. The proposed method is our first attempt for constructing a CWAC agent. We will continue develop other CWAC agents and integrating them into a multi-CWAC agent system.

Read the paper · More papers on PaperTik