Enhancing Decision Tree Performance through Stacking Ensemble Learning for Sentiment Analysis
Stanley Pratama Teguh, Jesslyn Trixie Edvilie, Ardhi Bagas Rangga Wardhana, Irene Anindaputri Iswanto, Setiawan Joddy · Procedia Computer Science · 2025
The natural language processing technique known as sentiment analysis deals with problems regarding how to achieve both high accuracy and easy interpretability levels. Decision Tree algorithms display clear interpretability patterns while showing diminished accuracy performance when operating independently for complex sentiment analysis operations. The study suggests ensemble learning stacking as an answer to boost Decision Tree performance while maintaining its clear interpretation features. The study draws its data from the IMDB movie review dataset while applying noise removal, lemmatization and TF-IDF feature extraction and subsequently implementing Synthetic Minority Oversampling (SMOTE). An evaluation of six candidate models including Random Forest, Extra Trees, SVM, KNN, Naïve Bayes, Logistic Regression was conducted together with Decision Trees to determine suitable combination algorithms for ensemble building. The stacking architecture brought together three base learners that included DT and the top candidate along with a distinct model while a meta-learner mirrored second base learner. The ensemble method that combined Decision Tree and Support Vector Machine achieved the best results with 88.08% accuracy, with precision and recall and F1-score matching that value at 88.08%. They outperformed exclusive Decision Tree with 73.46% accuracy. The addition of third-layer models including DT + SVM + Logistic Regression resulted in minimal accuracy improvement compared to simpler approaches (88.11%). The examined research shows that stacking produces effective results by combining different model strengths. SVM boundary optimization functions well in combination with DT hierarchical decisions to produce an effective pair. The addition of multiple layers resulted in performance improvement but caused unnecessary computing costs. The research demonstrates how stacking ensembles can produce powerful interpretable sentiment analysis results when combined with scenarios that demand both accuracy and transparency. Future research could combine deep learning techniques with new investigation methods to enhance interpretability and make systems more robust.