Hindi Text Summarization with Transformer-Based Ensembling and Refinement
Manisha Maharana, J Meghna, Ananya Kaushik, Anshika Ambavat, Amita Dev, Poonam Bansal, Aparna Kaushik · 2025
Text Summarization is one of the most important tasks of natural language processing (NLP) that revolves around the extraction of key text from a larger set of text and writing an informative summary. For Hindi and other low-resource languages, the unavailability of annotated datasets together with the complexity of the language renders the task extremely challenging. This work assesses the performance of two transformer models—mT5 and IndicBART—on abstractive summarization of Hindi text based on the SAR dataset, which is regarded as a standard benchmark for this type of task. The models are fine-tuned and tested separately based on ROUGE metrics. For further improving the quality of summarization, an ensemble strategy is adopted wherein the optimal summary is chosen among the two models per sample based on ROUGE-L recall. Experimental results demonstrate that the ensemble model outperforms the individual models, achieving ROUGE-1: 0.456, ROUGE-2: 0.27, and ROUGE-L: 0.416 (F1 scores)—showing better overall precision and recall than the other models. The results from this study prove that the ensemble of multilingual transformer models, alongside lightweight post-processing, improves summarization performance in low-resource languages, which leads to more efficient and reliable multilingual NLP systems.