A Review of Morphological Analysis Methods on Uyghur Language

Gvzelnur Imin, Mijit Ablimit, Askar Hamdulla · 2021

A morphological diverse language will form a huge collection of various types of words. As an agglutinative language, Uyghur language is composed of affixes connecting the front and back of the stem to form a large number of words. Based on the analysis of Uyghur language morphology, the first part focuses on the three language features of Uyghur language, including word formation and ambiguity, cohesion, and phonetic changes. The second part discusses the characteristics, applications and research purposes of Uyghur language morphological analysis. The third part introduces and describes in detail the methods, advantages and disadvantages and their characteristics based on the domestic and foreign research of morphological analysis. The fourth part respectively introduces several Uyghur language stemming methods and corresponding implementation cases, and embodies the characteristics of Uyghur language. The fifth part introduces the concatenated embedding of word embedding and character-level embedding to extract Uyghur language stems through BiLSTM-CRF model. First, obtain the word embedding of each word through the unlabeled Uyghur language corpus. Secondly, obtain the character feature embedding and then directly concatenate to obtain the concatenated embedding representation. Finally, the BiLSTM-CRF model is used to extract the stem of Uyghur language, and the accuracy rate reaches 89.21%. The conclusion part summarizes the word stem extraction, and looks forward to the future development trend of Uyghur language morphological analysis research, and also discusses the development direction of agglutinative language information processing.

Read the paper · More papers on PaperTik