The Identification and Classification of Unknown Words in Chinese : A N-Grams-Based Approach

Mei-Chu Wang, Chu‐Ren Huang, Keh-Jiann Chen · Institutional Repositories DataBase (IRDB) · 1995

In this paper, we propose a new approach to identify unknown words in Chinese. This approach adopts an n-grams program to sort out the collocating word / character sequences which are possible words and phrases in Chinese. In addition to proposing the criteria for identifying Chinese new words, was also classify these new words according to their structural and semantic characteristics. The corpus-based approach in identifying Chinese disyllabic words based on mutual information was first studied by Sproat and Shih [1]. The attempt here is to identify Chinese unknown words by collocations. Collocations are sequences of words that tend to appear together. In this paper we describe a set of techniques based on statistical methods for retrieving and identifying unknown words from a Chinese corpus. Here unknown words refer to words that are not included in the 90,000 entries CKIP Electronic Dictionary developed at Institute of Information Science, Academia Sinica. The results retrieved by the n-grams program will be crucial information for updating dictionaries. The n-grams program locates words in context and makes statistical observations to identify collocations. It produces a wide range of collocations which can be further sub-classified as abbreviational words, derived words, proper names, new words, ambiguous words, and collocating strings. The effectiveness of the n-grams program as a retrieval tool for unknown words is measured and evaluated.

Read the paper · More papers on PaperTik