A Review of the State-of-the-Art Techniques and Analysis of Transformers for Bengali Text Summarization

Iftekharul Mobin, Mahamodul Hasan Mahadi, Al‐Sakib Khan Pathan, A. F. M. Suaib Akhter · Big Data and Cognitive Computing · 2025

Text summarization is a complex and essential task in natural language processing (NLP) research, focused on extracting the most important information from a document. This study focuses on the Extractive and Abstractive approaches of Bengali Text Summarization (BTS). With the breakthrough advancements in deep learning, summarization is no longer a major challenge for English, given the availability of extensive resources dedicated to this global language. However, the Bengali language remains underexplored. Hence, in this work, a comprehensive review has been conducted on BTS research from 2007 to 2023, analyzing trends, datasets, preprocessing techniques, methodologies, evaluations, and challenges. Leveraging 106 journal and conference papers, this review offers insights into emerging topics and trends in Bengali Abstractive summarization. The review has been augmented with experiments using transformer models from Hugging Face and publicly available datasets to assess the Rouge score accuracy for Abstractive summarization. The extensive literature review conducted in this study reveals that before the advent of transformers, LSTM (Long Short-Term Memory) models were the dominant deep learning approach for text summarization across various languages. For transformers, one of the key datasets utilized was XL-SUM with the MT5 model emerging as the best performer among various contemporary multilingual models. These findings contribute to understanding the contemporary techniques and challenges in BTS. Furthermore, recommendations are made to guide future research endeavors, aiming to provide valuable insights and directions for researchers in this field.

Read the paper · More papers on PaperTik