Improving Sentiment Analysis Performance on Imbalanced Dataset Using Data Resampling and Statistical Feature Selection

Muhammad Fachrie, Aina Musdholifah, Sri T. Hartati · 2024

Imbalanced dataset is one of major challenges in developing machine learning model. This imbalance problem leads to bias in the classification model that ends in low classification performance. In many cases of sentiment analysis, the imbalance problem often arises in the class distribution, where opinions tend to be either ‘positive’ or ‘negative’ depending on the concerning topic. Besides, the presence of high-class overlap within the imbalanced dataset also has a negative impact on the classification performance. Furthermore, the raw text data is usually transformed using the well-known TF-IDF feature extraction. This transformation generates a set of high-dimensional features containing the statistical value of terms extracted from the raw text data that makes it very complex to classify. Therefore, this work proposes the use of data resampling and statistical feature selection to improve the classification performance of sentiment analysis on high-dimensional datasets and imbalanced class distribution. As a result, statistical feature selection combined with data resampling methods successfully improves the classification performance of five different machine learning algorithms on two imbalanced datasets. The proposed approach achieves better accuracy while reducing the dataset dimensionality by up to 73% and 89% of the original features on two different datasets.

Read the paper · More papers on PaperTik