Part-of-Speech Identification for Unknown Chinese Words Based on k-Nearest Neighbors Strategy

Sun Mao · Chinese Journal of Computers · 2000

Unknown word processing plays an important role in many natural language application systems aiming at large scale unrestricted texts. The task of part of speech identification is to automatically assign a part of speech tag to an unknown word with empty part of speech information. A part of speech identification algorithm based on k- nearest neighbors strategy is presented in this paper. The preliminary experiment, supported by a Chinese corpus of 100M characters and a part of speech annotated corpus of 0.6M characters, shows that the average accuracy rates of the algorithm can reach 99.21%, 84.73%, 70.67% for Chinese words of nouns, verbs and adjectives respectively.

Read the paper · More papers on PaperTik