A Morphology-Driven Approach to NLP for a Low-Resource, Highly Complex Language
Irakli Kardava · Vietnam Journal of Computer Science · 2025
This paper presents a study investigating the optimization of well-known NLP algorithms and approaches for the Georgian language, known for its unique linguistic features. Standard methods effective for well-resourced languages, including pretrained models like mBERT and embedding methods such as FastText, may lack flexibility and efficiency when applied to Georgian, often resulting in increased complexity and effort. To address these challenges, we propose a novel approach that leverages Georgian’s rich morphology, including case inflections, extensive suffixation, verb agreement, and conjugation patterns. This method refines algorithms such as Minimum Editing Distance, Text Classification, Language Modeling, and word-level semantic similarity by incorporating language-specific characteristics. Our approach reduces data sparsity and model complexity while preserving accuracy. Although developed for Georgian, it is also relevant for other fusional and agglutinative languages and contributes to reducing dependence on large corpora, supporting the creation of more human-like text.