Hybrid Sentiment Analysis Model combining BERT and Gradient Boosting for Enhanced Text Classification

Diksha Dhiman, V. Asha, Mithili Devi, Swasthik Shetty, Upendra Gurav · 2025

Sentiment analysis is an essential task of Natural Language Processing (NLP) because it allows our computer to automatically classify a piece of text into positive, negative, or neutral. As effective as these techniques are, typical machine learning family techniques often do not encapsulate the complex context relationships present in a given text. Transformer-based models like BERT, on the other hand, require significant computational power, while achieving a good level of semantic understanding. In order to resolve these difficulties, this paper proposes a hybrid sentiment analysis model that utilizes the classification strength and effectiveness of Gradient Boosting algorithms with the contextual feature extraction strengths of BERT. The Sentiment140 dataset was used to test and build the model. Preprocessing included data cleaning, tokenization, and balancing techniques. Contextual embeddings from BERT were passed to XGBoost, LightGBM, and CatBoost classifiers, with XGBoost yielding the best performance. Experimental results demonstrate that the hybrid model outperforms both standalone BERT and traditional Gradient Boosting models, achieving an accuracy of 92% and an F1-score of 0.92. This architecture presents a practical, efficient solution for real-world sentiment analysis, offering a balanced trade-off between accuracy and computational cost. Future work will focus on improving model explainability, enhancing multilingual capabilities, and enabling real-time deployment in resource-constrained environments.

Read the paper · More papers on PaperTik