n-Gram-based indexing for Korean text retrieval

Joon Ho Lee, Hyun Yang Cho, Hyouk Ro Park · Information Processing & Management · 1999

Two groups of indexing methods such as word-based indexing and morpheme-based indexing have been investigated in the literature of Korean text retrieval. The word-based indexing eliminates the suffix of a word, and generates its remaining stem as an index term. The index term is often a compound noun, which results in the serious decrease of retrieval effectiveness. The morpheme-based indexing overcomes the problem of compound nouns by decomposing a compound noun into simple nouns. It, however, requires a large dictionary and complex linguistic knowledge. In this paper we propose a new indexing method based on n-grams, which can handle compound nouns effectively without dictionaries and complex linguistic knowledge. We also evaluate the indexing method for Korean texts through experiments. Experimental results show that the n-gram-based indexing is considerably faster than the morpheme-based indexing, and also provides better retrieval effectiveness.

Read the paper · More papers on PaperTik