Generative Pre-Trained Transformer for Kazakh Text Generation Tasks

Gulmira Tolegen, Alymzhan Toleu, Rustam Mussabayev, Багашар Жумажанов, Gulzat Z. Ziyatbekova · 2023

This paper presents an empirical study evaluating text generation models for the low-resource and morphologically complex Kazakh language. In this study, we leveraged a transformer-based neural architecture. Initially, we trained a large language model for the Kazakh language using a substantial text corpus. Then, we fine-tuned the base model using a specialized question-answering dataset. The proposed pretrained language model (PLM) was evaluated on a conversation response generation task, investigating its performance across domains and languages. The findings of cross-lingual comparison underlines the challenges of achieving good results for low-resource languages like Kazakh with current solutions. The experimental findings revealed that despite the low-resource characteristics of Kazakh language, Kazakh's PLMs and its fin-tuned model for question answering were obtained and its performance for the BLEU score was 8.5%, which was better than the random guess 4.01%. Additionally, Two Kazakh text corpora were collected for this work. One of datasets comprised a large collection of Kazakh text from different domains and was used to train a large language model. Second datasets were specifically collected for the question-answering task for Kazakh and Russian languages.

Read the paper · More papers on PaperTik