Automated Root Identification and Pattern Detection for Agglutinative Language Processing
Irakli Kardava · 2024
Achieving high accuracy in various Language Models (LMs) and Large Language Models (LLMs) relies heavily on extensive training datasets. However, as natural languages evolve, words change form, and new words emerge, necessitating the continuous updating of these models. Incorporating new languages also adds to this complexity. Since texts are composed of words, these updates hinge on changes in word forms, making words and their forms the most critical components of the training set. This research focuses on generating grammatically correct word forms for different parts of speech using natural language processing (NLP) techniques, eliminating the need for manual pre-recording of the data and its storage in databases. This is crucial when dealing with agglutinative languages, and is especially valuable in the case of Less Resourced Languages (LRLs). Our experiments and findings primarily revolve around the Georgian language, but we indicate that this approach can be effectively applied to other agglutinative languages as well.