Resolving the noun sense ambiguity of Korean based on concept co-occurrence information
Namwon Heo, Huifeng Li, Kyonghi Moon, Jong-Hyeok Lee, Geunbae Lee · Open Access System for Information Sharing (Pohang University of Science and Technology) · 1999
Unlike verbal words, which are usually accompanied by a variety of linguistic knowledge, nominal words suffer from lack of such linguistic clues to identify their meaning in a context. This paper proposes a corpus-based word sense disambiguation (WSD) method, especially focusing on nominal words. Most previous research has restricted the use of linguistic knowledge to the lexical (i.e. surface word) level. On the contrary, we rely on concept co-occurrence information extracted from a sense-tagged corpus. The sense-tagged corpus is automatically constructed from a Japanese raw corpus by an existing Japanese-to-Korean MT system with high translation quality. The concept co-occurrence information consists of two parts: local collocation patterns (LCPs) and unordered sets of surrounding words (USWs) encoded with the Kadokawa thesaurus. To improve the recall rate and also to save storage space while keeping a high precision, the extracted conceptual information is type-abstracted into higher levels. The WSD algorithm is performed at four levels, each level using a different knowledge like selection restrictions of verbs, concept co-occurrence information (LCPs and USWs) of nouns, and heuristics. In an experiment, it achieved the average precision of 82.4%, which is an improvement of the baseline by 14.6%. Considering that the test corpus is completely irrelevant to the learning corpus, this turns out that the proposed method may be very effective in WSD.