Lemmatization of big data in the Kazakh language

Diana Rakhimova, Aliya Turganbayeva · 2019

In paper are considered existing algorithms for automatically isolating the bases for a number of natural languages and possible ways of synthesizing a normal form of a word for the Kazakh language. In paper are described the complete system of endings of the Kazakh language. The article presents the classification of affixes Kazakh language. The paper proposes a new approach to constructing a lemmatization algorithm for the Kazakh language on the basis of a complete set of endings of the Kazakh language. The lemmatization algorithm will be used for Kazakh language information retrieval to finding specific words in the documents by stemming base of word.

Read the paper · More papers on PaperTik