Extractive Question Answering for Kazakh Language

Magzhan Shymbayev, Yermek Alimzhanov · 2023

This article provides research and development of an extractive question answering system based on the BERT-like model for the Kazakh language. Developing an extractive question answering system requires large training datasets - tens of thousands of annotated question-answer pairs. Such datasets are not available in the majority of languages, including Kazakh. To address this issue, the Kazakh Question Answering Dataset (KazQA) is introduced, which is based on the Stanford Question Answering Dataset (SQuAD) and generated through machine translation using the Google Cloud Translation API. Different large pretrained contextual language models are used as the baseline models - ALBERT and multilingual BERT and are compared with the newly trained monolingual Kazakh model KazBERT. The results demonstrate that the proposed approach can effectively generate question answering systems in low-resourced Kazakh language.

Read the paper · More papers on PaperTik