Enhancing Lemmatization for Mongolian and its Application to Statistical Machine Translation
Odbayar Chimeddorj, Atsushi Fujii · International Conference on Computational Linguistics · 2012
Lemmatization is crucial in natural language processing and information retrieval especially for highly inflected languages, such as Finnish and Mongolian. The state-of-the-art method of lemmatization for Mongolian does not need a noun dictionary and is scalable, but errors of this method are mainly caused by problems related to part of speech (POS) information. To resolve this problem, we integrate POS tagging and lemmatization for Mongolian. We evaluate the effectiveness of our method and its contribution to statistical machine translation.