Boosting ChatGPT Sentiment Classification with Stacking and Majority Voting on Skewed Data

Dharmaraj Rajaram Patil · Vietnam Journal of Computer Science · 2025

In the field of natural language processing (NLP), sentiment analysis is very important for comprehending user views and opinions. By using the Kaggle ChatGPT sentiment analysis dataset, this work investigates the use of sophisticated ensemble machine learning algorithms to enhance ChatGPT’s sentiment classification performance. Due to its skewed distribution, the dataset poses particular difficulties in reaching high performance metrics. For sentiment classification, we evaluate the performance of many boosting techniques such as AdaBoost, Gradient Boosting Machine (GBM), XGBoost, LightGBM (LGBM), and CatBoost. Moreover, ensemble techniques such as majority voting and stacking are used to improve classification results. While majority voting combines predictions for robust classification, stacking specifically uses many base learners in conjunction with a meta-model to maximize performance. With a classification accuracy of 88.57%, precision of 88.67%, recall of 88.57%, and an F-measure of 88.57%, our tests show that stacking performs better than any other technique. Its efficient mistake mitigation is further demonstrated by its False Positive Rate (FPR) of 0.09 and False Negative Rate (FNR) of 0.13. These findings highlight how ensemble methods — in particular, stacking — can be used to handle skewed data and enhance sentiment classification capabilities. In light of ChatGPT and related conversational artificial intelligence (AI) technologies, the results offer important new information for the creation of trustworthy sentiment analysis systems.

Read the paper · More papers on PaperTik