Improving Classification Accuracy With Preprocessing Techniques For Sentiment Analysis

Ayu Sekar Safitri, Inung Wijayanto, Sugondo Hadiyoso · 2024

Sentiment is an opinion expressed by individuals, ranging from positive to negative meanings. Sentiment analysis can provide valuable information, criticism, or advice. With the growing number of social media users on platforms like Twitter, now known as X, people can express themselves more freely. However, processing this vast amount of data for sentiment analysis requires a preprocessing stage to clean the data, ensuring the system captures the most relevant words for analysis. This study focuses on the preprocessing stage, using lemmatization for data cleaning, the Synthetic Minority Over-sampling Technique (SMOTE) to address class imbalance, and Term-Frequency-Inverse Document Frequency (TF-IDF) for feature extraction. The Twitter US Airline Sentiment dataset was utilized for this study. Post-preprocessing, the dataset was classified using machine learning techniques such as K-Nearest Neighbor (K-NN), AdaBoost, Decision Tree, Random Forest, Support Vector Machine (SVM), and Gaussian Naive Bayes. The proposed approach significantly improved classification results, with F1-Scores of 94.9%, 96.21%, 95.55%, 68.99%, 83.24%, 90%, and 52.75% for each method, representing a 7% to 31% increase in precision, recall, and F1-score compared to the baseline without lemmatizatiom, SMOTE and TF-IDF. These results highlight the great potential of combining preprocessing techniques with various machine learning algorithms to extract diverse sentiments from social media data, enabling more detailed analysis and a greater understanding of public opinion. This study underscores the importance of selecting appropriate preprocessing techniques to enhance the accuracy and effectiveness of sentiment analysis.

Read the paper · More papers on PaperTik