Exploring State-of-the-Art LLMs from BERT to XLNet: A study over Question Answering

L. M. Parra-Navarro, Evelyn C. S. Batista, Marco Aurélio C. Pacheco · 2024

In recent years, both academia and industry have witnessed significant advancements in Large Language Models (LLMs) research, with models like ChatGPT garnering extensive attention from society. These advancements in LLM technology have exerted a profound influence on the entire AI community, potentially revolutionizing how we design and utilize AI systems. Among Natural Language Processing (NLP) tasks, Question Answering (QA) has gained increasing attention. In this paper, we provide an overview of the advancements in LLMs, covering background, major findings, and fine-tuning experiments conducted from BERT to XLNet models on the SQuAD v1.1 and SQuAD v2.0 datasets for the QA task. Our evaluation of various encoder-only models on SQuAD tasks reveals that RoBERTa consistently demonstrates the best performance, achieving the highest Exact Match (EM) and F1 scores on both the SQuAD 1.1 and SQuAD 2.0 datasets. Additionally, we find that Flan T5-base yields even better results, boasting an EM of 77.9% and an F1 Score of 81.2%.

Read the paper · More papers on PaperTik