Investigating the Effectiveness of Feature Extraction Techniques in Predicting Emotions from Indonesian Tweets Using Machine Learning
Jeremy Christopher N. S., Julius Lie, Ghinaa Zain Nabiilah, Jurike V. Moniaga · 2024
The rapid growth of social media platforms, particularly Twitter, resulted in an unprecedented volume of user-generated content, making it essential to understand the sentiments and emotions expressed in these texts. This study focuses on enhancing sentiment analysis techniques for Indonesian tweets, an area of growing importance given the increasing number of Indonesian Twitter users. The primary objective is to improve the accuracy of sentiment analysis by comparing the performance of two machine learning algorithms: Support Vector Machine (SVM) and Naïve Bayes, using two text representation techniques: TF-IDF and Count Vectorization. A dataset of 4,403 Indonesian tweets, labeled with five different emotion classes-love, anger, sadness, happiness, and fear that underwent preprocessing to ensure uniformity. The analysis revealed that the SVM achieved an accuracy of 65% using TF-IDF, outperforming Naïve Bayes' 60% accuracy. Conversely, with Count Vectorization, Naïve Bayes outperformed SVM, achieving an accuracy of 66% compared to SVM's 61 %. The results indicate that while SVM excels with TF-IDF, Naïve Bayes performs better with Count Vectorization. This research highlights the necessity of tailoring sentiment analysis approaches to the specific characteristics of the dataset and contributes to the advancement of sentiment analysis methodologies for Indonesian tweets and other social media data.