Comparative Analysis of Large Language Models Fine-tuned on a Singular Dataset

Adithya Tadaga Gangadhar, Harsh Pranav Rao, Shylaja Ss · 2024

This paper presents a comparative analysis of several state of the art language models fine-tuned for question answering tasks using the IEEE research papers dataset. The models evaluated include GPT-2, BERT, BART, Llama Mistral and RAG with GPT-3.5-Turbo. Each model was fine tuned and evaluated based on their accurate responses generated to the questions posed on different categories of academic paper titles and abstracts. The study employs evaluation metrics such as BLEU score, ROUGE 1 and ROUGE 2 to assess linguistic quality, factual accuracy and content overlap between generated answers and reference texts. The results obtained show the different performance characteristics of the models, providing information into their suitability for question answering in academic contexts. This paper aims to inform researchers and other users to select the appropriate models based on specific task requirements, improving the application of natural language processing in academic research domains.

Read the paper · More papers on PaperTik