A Machine Learning Approach to Korean Language Stemming
Se-hyeong Cho · Journal of Korean institute of intelligent systems · 2001
Morphological analysis and POS tagging require a dictionary for the language at hand. In this fashion, though, it is impossible to analyze a language without a dictionary. We also have difficulty if significant portion of the vocabulary is new or unknown. This paper explores the possibility of learning morphology of an agglutinative language, in particular, Korean language, without any prior lexical knowledge of the language. We use unsupervised learning, in that there is no instructor to guide the outcome of the learner, nor any tagged corpus. Here are the main characteristics of the approach: First, we use only raw corpus without any tags attached or any dictionary. Second, unlike many heuristics that are theoretically ungrounded, this method is based on statistical methods, which are widely accepted. The method is currently applied only to Korean language, but since it is essentially language-neutral, it can easily be adapted to other agglutinative languages.