Comparative Analysis and Application of Large Language Models on FAQ Chatbots
Feriel Khennouche, Youssef Elmir, Nabil Djebari, Larbi Boubchir, Abdelkader Laouid, Ahcène Bounceur · 2024
This paper presents a comprehensive evaluation of advanced large language models (LLMs) including GPT, BART, BERT, and T5, fine-tuned using a specialized FAQ dataset from ESTIN, a higher education institution. The study aims to assess the performance of these LLM models in generating accurate and contextually relevant responses to common queries. These models were evaluated using key metrics such as evaluation loss (eval loss) and ROUGE scores (ROUGE-1 and ROUGE-2), which measure the alignment between the generated responses and the reference answers. The experiment results show that BERT outperforms the other models, achieving a lowest eval loss and highest ROUGE-1 and ROUGE-2 scores, making it the most effective model in this context. The findings underscore the potential of BERT in educational AI applications. Finally, some future research directions are discussed, including the integration of additional models and the enhancement of the FAQ dataset to further improve performance.