Extractive Indonesian News Text Summarization using DistilBERT, IndoBERT, MBERT, and RoBERTa
Joseph Vincent Liem, Michael Ivan Santoso, Stevan Pohan, Muhammad Fikri Hasani, Ayu Maulina · Procedia Computer Science · 2025
Reading has become an essential activity in today’s digital era, where vast amounts of textual information are generated daily across various online platforms. Text summarization has become an important solution that allows users to easily understand the main points of the information. As a result, the field of Natural Language Processing (NLP) has experienced significant growth, especially with the emergence of transformer-based models that have pioneered new advancements. While extractive summarization using BERT-based transformer models has shown promising results in various languages, its application to Indonesian language content remains limited and understudied. This study evaluates the effectiveness of transformer-based models DistilBERT, IndoBERT, mBERT, and RoBERTa for extractive summarization of Indonesian news articles using the Liputan6 dataset. Evaluation was conducted using the ROUGE(Recall-Oriented Understudy for Gisting Evaluation) metric, focusing on ROUGE-1, ROUGE-2, and ROUGE-L to assess each model’s capability in capturing unigrams, bigrams, and long sequence matches. IndoBERT, trained exclusively on Indonesian text, achieves the best performance across all ROUGE metrics. ROUGE- 1 metric with a score of 0.4006, ROUGE-2 metric with a score of 0.2936, and ROUGE-L with a score of 0.3905. This demonstrates its superior capability in understanding and summarizing Indonesian content compared to mBERT and DistilBERT, which performs slightly worse despite supporting multiple languages. RoBERTa shows lower performance, likely due to their primary training on English text.