Evaluating Large Language Models for Sentence Augmentation in Low-Resource Languages: A Case Study on Kazakh
Zhamilya Bimagambetova, Dauren Rakhymzhanov, Assel Jaxylykova, Alexander Pak · 2023
Large language models (LLMs) have revolutionized natural language processing (NLP) and demonstrated exceptional performance in various NLP tasks for widely spoken languages. However, their efficacy in handling low-resource languages remains an area of concern. This study investigates the performance of LLMs, particularly GPT-3, in sentence augmentation tasks for a low-resource language, Kazakh. We employ a blind peer review methodology, where five native Kazakh annotators assess the quality of LLM-generated augmentations. The results reveal that LLMs excel in popular languages like English, Chinese, and German, but face challenges with low-resource languages due to limited training data. This work sheds light on the importance of improving LLMs' adaptability and relevance to address the unique needs of low-resource languages. Further research could enhance the augmentation capabilities of LLMs in scenarios with limited data sources, ensuring their effectiveness in promoting linguistic diversity and inclusivity. Furthermore, this study underscores the significance of cross-language transfer learning and data collection efforts to empower LLMs in supporting linguistic diversity and fostering inclusivity across the global language landscape.