Sentiment Analysis of Indonesian News Texts Using IndoBERT and IndoRoBERTa
Nasywaa Rihaadatul'Aisy, Endang Wahyu Pamungkas · 2025
This study investigates the performance of transformer-based models-IndoBERT and In-doRoBERTa-alongside a Support Vector Machine (SVM) baseline for sentiment analysis of Indonesian news texts. Using a dataset of 3.538 articles, the research employed a systematic pipeline comprising data collection, labeling, preprocessing, undersampling to address class imbalance, modeling, cross-validation, and deployment. IndoBERT achieved the highest accuracy of 70.76 % (F1-score: 0.71), validated by a 5-fold cross-validation mean accuracy of 70.37 %, followed by IndoRoBERTa at 68.79 % (F1-score: 0.69). The SVM, optimized via GridSearchCV, attained 61 % accuracy (F1-score: 0.57). Key challenges included the limited dataset size, class imbalance (45 % neutral, 33 % negative, 22 % positive), and the complexity of Indonesian linguistic nuances, which hindered positive sentiment detection-particularly for SVM. The study recommends expanding the dataset to over 5.000 articles, exploring alternative imbalance-handling techniques such as SMOTE, and investigating ensemble methods or lighter models like DistilBERT.