Turkish Question - Answer Dataset Evaluated with Deep Learning
Kadir Tutar, Olcay Taner Yıldız · 2024
The field of Natural Language Processing (NLP) poses various challenges, one of which is the task of question-answering which involves comprehension, multiple-choice, yes-no questions, and more. Unfortunately, creating question-answer dataset is one of the main obstacles to these studies, especially in some languages such as Turkish where dataset availability is limited. To address this scarcity, this study developed a Turkish question-answering dataset from the comprehensive English SQuAD dataset using Microsoft Translator. We initially checked all translator output and observed that we can use approximately 60% of data after updating index of answer. Then 16% of all data was refined in Turkish language for increasing data amount. As a result, we successfully created 78% of the SQuAD in Turkish language. Additionally, a question-answering deep learning model was built to test the dataset's usability and model performance was assessed using various metrics, domain-specific considerations, translation quality, and linguistic statistics.