Multilingual Stemming and Term extraction for Uyghur, Kazak and Kirghiz
Mijit Ablimit, Sardar Parhat, Askar Hamdulla, Thomas Fang Zheng · 2018
Stemming and term extraction is an important and difficult step on NLP for low resource languages. Inflectional structure and noisy data aggravate this problem for less popular agglutinative languages. A morphological analyzer with the longer context can utilize resources and provide reliable semantic and syntactic information, and effectively reduce ambiguity on stemming and term detection. There are some previous works on stemming on Uyghur texts based on simple morphology like affix and manually collected rules. But the limited information of smaller context and lower segmentation accuracy cost the reliability. We developed a sentence level multilingual morphological processing tool for Uyghur, Kazak, and Kirghiz languages. This tool can provide sentence level morpheme extraction with 98% accuracy, and further analysis like word embedding and longer context modelling can extract remaining unseen and infrequent stems reliably. Combined with word embedding this tool provides a more reliable way of term extraction from large number of noisy text available from internet.