COMPARATIVE ANALYSIS OF LARGE LANGUAGE MODELS FOR THE APPLICATION OF SCIENTIFIC ARTICLE SUMMARIZATION

Rishabh Saxena, Shubhangi Singh, Preeti Dubey · Procedia Computer Science · 2025

Automatic text summarization has become an essential tool for academics to stay up to date with the newest advancements due to the high rise of scientific publications. Abstractive text summarisation tasks can be done using LLMs. However current summarization methods often struggle to capture the nuanced and technical details present in research papers and there are prevalent research gaps in the realm of abstractive summarisation using generative AI. This study aims to analyse the efficiency of LLM models – GPT-3.5, LLaMa 3, Mixtral 8x7b and Gemma-2, in condensing detailed scientific literature into manageable summaries, and their further evaluation based on ROUGE, BERT and BLEU scores. The results show that GPT-3.5 and LLaMa 3 generate more coherent and contextually accurate summaries with a BERT score of 0.21 and 0.19 respectively, even though Mixtral 8x7B performs exceptionally well in quantitative measurements. Despite its coherence, Gemma-2 receives poorer performance in both qualitative and quantitative assessments. These findings bring out the value of integrating quantitative and qualitative evaluations to gain a thorough understanding of summarization.

Read the paper · More papers on PaperTik