Machine Reading Comprehension with Tamil Short Story MRC Dataset

A. V. Ann Sinthusha, E.Y.A. Charles, Ruvan Weerasinghe · 2024

Multilingual transformer models are considered to be a solution for Natural Language Processing (NLP) in under-resourced languages like Tamil. Machine Reading Comprehension for the Tamil language was an under-studied task compared to other Tamil natural language processing tasks and other languages. This is mainly due to the lack in the presence of language-specific pre-trained transformer models and machine reading comprehension datasets for the Tamil language. A quality dataset is an important ingredient for all kinds of natural language processing tasks. To analyse the performance of the machine reading comprehension task for the Tamil language an MRC dataset based on Tamil short stories was created as part of this research. A question count-based analysis was performed to verify the variety of questions. This analysis found that the Tamil Short Story for MRC task dataset possesses a similar ratio of question types compared to SQuAD 1.1. Further, to verify the effect of utilising language-specific resources, this dataset and the famous SQuAD 1.1 dataset of the English language have been used to analyse the performance of the transformer models, namely, multilingual DistilBERT, multilingual BERT, XLM-RoBERTa, MuRIL, and RemBERT. These models were selected due to their support for the Tamil language. As per the analysis, the RemBERT and MuRIL models were able to produce the highest BERT Score of 85.33% confirming the contribution of language-specific resources.

Read the paper · More papers on PaperTik