Machine Comprehension System in Tamil and English based on BERT
G. Srivatsun, S. Thivaharan, Bharath Kumaar K S, S Sudharsan · 2022 3rd International Conference on Electronics and Sustainable Communication Systems (ICESC) · 2022
Question Answering is a part of information retrieval. It is a task of machine-reading comprehension and presents a medium to assess the ability of machines to understand human language. A question-and-answer system is preferable to search engines that return a collection of documents. Any question answering system's high level design consists of three primary components: Question analysis, retrieval of context, answer is being extracted. With an availability of training resources and Transformer-based models trained on huge English corpora, the precision and accuracy of English Question Answering systems has increased dramatically over the years. Such big datasets, however, are not available for low-resource languages like Tamil. For low-resource languages, multilingual BERT (mBERT) models are utilized. The translations from the same language family is used to supplement the available data and fine-tune the mBERT-based model. Cross-lingual learning uses zero, and few-shot techniques applied to transfer the knowledge of a QA model trained on many source examples to a given target language with fewer training data. Pre-training and fine-tuning of the model improves performance, according to model tests.