Development and Evaluation of a Small Kazakh Language Corpus to Improve the Efficiency of Multilingual NLP Systems in Low-Resource Environments
Arailym Tleubayeva, Sultan Aubakirov, Aisultan Tabuldin, Shomanov Aday · 2025
This study tackles NLP challenges in low-resource settings by developing the Small Kazakh Language Corpus—a high-quality, annotated collection of Kazakh texts sourced from news, scientific publications, and Wikipedia. The corpus was used to fine-tune two models for masked language modeling: the multilingual XLM-RoBERTa-base and the Kazakh-specific nur-dev/roberta-kaz-large. Fine-tuning notably improved XLM-RoBERTa-base's accuracy from 48.84% to 68.85% and F1-score from 41.96% to 68.62%, while nur-dev/roberta-kaz-large showed more modest gains. These findings demonstrate the critical role of language-specific resources in enhancing multilingual NLP systems and provide a solid foundation for further research in applications such as machine translation, sentiment analysis, and question answering.