Algorithm Analysis of Decision Tree, Gradient Boosting Decision Tree, and Random Forest for Classification (Case Study: West Java House of Representatives Election 2019)
Irena Irmalasari, Latifa Dwiyanti · 2023
In an increasingly modern era, machine learning-based technologies offer solutions to everyday problems. Machine learning enables us to build models that process data and train them to provide predictive results. We call predictions with categorical labels ‘classifications’. We commonly use several basic models for classification, including Decision Trees, Support Vector Machines, Linear Regression, and so on. We can combine these basic models to create an ensemble model, consisting of a boosting algorithm and a bagging model. We will test these two types of algorithms on the profile data of the candidates for the HoR West Java in the 2019 legislative elections. In this paper, we compare the Random Forest bagging algorithm and the Gradient Boosting Decision Tree boosting algorithm. The Decision Tree algorithm is also compared as a baseline. Prior to the comparison, we conduct data pre-processing and data exploration to identify the data features that influence the winning label for feature selection. Furthermore, we address the data imbalance by using Borderline-SMOTE before training the model. We compare the test results based on the F1 score. Our findings reveal that the use of selected features and the Borderline-SMOTE1 sampler significantly improved the F1 score. Additionally, the best-performing model for predicting victory with the given data is the Gradient Boosting Decision Tree, with parameter tuning and using feature selection results, which achieved an F1 score of 0.7445