Comparing Random Forest and XGBoost for Sentiment Classification of Student Social Media Posts: A Case Study on Pre-Board Exam Stress
Sumit Saklani, Mahesh Manchanda, Geetanjali Bisht, Kalyan Thapa · 2025
With the rise of social media, students have found a new way of expressing their emotions during high-stress situations such as a pre-board exam. These platforms create an enormous amount of textual information from which we can analyze a students' stress level and sentiment. This research analyzes how well two ensemble machine learning approaches, Random Forest (RF) and XGBoost, perform in classifying students’ sentiments based on social media posts. The dataset analyzed in this study consists of anonymized social media posts during the pre-board exam and was created specifically using a new preprocessing pipeline that helps in understanding the complexities posed by short, noisy and informal text. The pipeline included state of the art slang normalization using custom dictionaries and Term Frequency-Inverse Document Frequency (TF-IDF) for feature extraction. The models were extensively tested on the metrics of accuracy, precision, recall and F1 score. Results showed that while robust performance was made by RF, XGBoost outdid it across all metrics with accuracy of 88.9% alongside stellar performance in precision and recall. This indicates how well XGBoost generalizes and computes efficiently especially with class imbalance. The research not only explain the significance of pre-processing during a sentiment analysis but also helps performers polish their skills over time.