Comparative Performance of Machine Learning Algorithms for Sentiment Analysis: The Role of PCA as a Dimension Reduction Technique

Burak Borhan, Yasin Ortakcı, Amit Lathigara · 2025

This study investigates the performance of various machine learning algorithms in the domain of sentiment analysis. It also uses a set of pre-processing techniques including tokenization, lemmatization, scaling, and Principal Component Analysis (PCA). Text review data from Amazon, IMDb, and Yelp is evaluated, and Word2Vec-based feature representations are used. Our experiments show that Support Vector Machine (SVM) often emerges as the best contender, achieving commendable results, especially when combined with PCA. However, we also observed cases where Logistic Regression or Random Forest algorithms performed comparable to SVM under certain data distributions. This finding suggests that no single algorithm is universally dominant in all scenarios. Meanwhile, the findings for real-time classification highlight that effective sentiment analysis depends not only on choosing a robust classifier, but also on using the right mix of pre-processing and dimension reduction techniques. While the combination of SVM and PCA often stands out, our results highlight the importance of adapting the approach to the dataset at hand and fine-tuning the hyper parameters. The comprehensive comparison provided in this work provides insights for researchers and practitioners aiming to optimise sentiment analysis workflows, enabling more accurate identification of positive and negative sentiments across diverse textual corpora.

Read the paper · More papers on PaperTik