Unknown word identification for Chinese morphological analysis
Chooi Ling Goh · Institutional Repositories DataBase (IRDB) · 2006
Since written Chinese does not use blank spaces to indicate word boundaries, segmenting Chinese texts becomes an essential task for Chinese language processing.Besides word segmentation, we also need to identify the part-of-speech (POS) tags of the words.The segmentation and POS tagging process are denoted as morphological analysis.During the process of word segmentation, two main problems occur: segmentation ambiguities and unknown word occurrences.There are basically two types of segmentation ambiguities: covering ambiguity and overlapping ambiguity.These ambiguities are dealt with known words.For the unknown word problem, we need to detect them from the text based on the context.In this report, we have focused on the problem of unknown words and proposed some machine-learning based methods towards solving it.Besides, we also face the ambiguity problem with POS tagging because a single word can hold multiple POS tags and it depends on the context to decide which one is the correct answer.Furthermore, if the word is unknown, then we need to guess the POS tag based on the word components and contexts.At the end of the research, we have built a practical morphological analyzer which can be freely used by anyone for research purpose.In order to build a practical system, a reasonable size dictionary is needed.The initial dictionary is built from the Penn Chinese Treebank corpus v4.0 and contains only 33,438 entries.Since the initial dictionary is quite small, the unknown word detection method is applied to huge raw texts in order to extract new words to be added into the system dictionary.We have successfully constructed a dictionary with 120,769 entries.Finally, we propose a two-layer morphological analysis to cater for two sets of outputs.The first layer produces the minimal segmentation unit