Synthetic Dataset Creation and Fine-Tuning of Transformer Models for Question Answering in Serbian

Aleksa Cvetanović, Predrag Tadić · 2023

In this paper, we focus on generating a synthetic QA dataset using an adapted Translate-Align-Retrieve method. We created the largest Serbian QA dataset, which we name SQuAD-sr. To acknowledge the script duality in Serbian, we generated both Cyrillic and Latin versions of the dataset. We investigate the dataset quality and use it to fine-tune several pre-trained models. Best results were obtained by fine-tuning the BERTić model on Latin SQuAD-sr dataset, achieving 73.91% Exact Match and 82.97% F1 Score on the benchmark XQuAD dataset, which we translated into Serbian for the purpose of evaluation. The results show that our model exceeds zero-shot baselines, but fails to go beyond human performance. We also note the advantage of using a monolingual pre-trained model over multilingual, as well as the performance increase gained by using Latin over Cyrillic. Finally, we conclude that SQuAD-sr is of sufficient quality for fine-tuning a Serbian QA model, and can be used in absence of a manually crafted dataset.

Read the paper · More papers on PaperTik