Optimizing IndoRoBERTa Model for Multi-Class Classification of Sentiment & Emotion on Indonesian Twitter

Yogie Oktavianus Sihombing, Reza Fuad Rachmadi, Surya Sumpeno, Moh. Jabir Mubarok · 2024

Text classification techniques using BERT have shown promising results on various natural language processing problems, especially sentiment analysis over Twitter. In this study, the IndoRoBERTa model is optimized to perform sentiment and emotion classification modeling on datasets sourced from Indonesian Twitter. IndoRoBERTa is a modified BERT called RoBERTa which is trained in Indonesian. In the comparison test, the proposed IndoRoBERTa model performed better than other Indonesian BERT variants. The accuracy and F1-score values for the proposed sentiment classification model resulted in 98% accuracy and 97.4% F1-score. In the meantime, the suggested emotion classification model yielded an 83% F1-score and an accuracy of 82.7%. Based on the confusion matrix, the IndoRoBERTa model for sentiment classification performed very well in all classes. Meanwhile, the IndoRoBERTA model for emotion classification needs further evaluation to understand why it is often incorrectly predicted and how to improve the accuracy of the classes. During the classification test, several raw tweet samples were used, the high probability values for each classification show that the model is confident in determining the sentiment and emotion of each tweet. These results show great potential for exploring users’ emotional responses in more detail.

Read the paper · More papers on PaperTik