Handling Missing and Imbalanced Data to Improve Generalization Performance of Machine Learning Classifier
Andri Aryarasyid Dharmasaputro, Nadhif Muhammad Fauzan, Meta Kallista, Ig. Prasetya Dwi Wibawa, Purba Daru Kusuma · 2022
Exploratory data analysis (EDA) is an important process for creating a machine learning model. Through data preprocessing, we can see the characteristics of the data and how to handle them. Data preprocessing is part of EDA that handles problems in the data before creating training and testing datasets. As a study case, this research uses the air pollution dataset published by the Ministry of Environment and Forestry of the Republic of Indonesia. The dataset has missing value and imbalanced class problems in which the previous research used complete case analysis for missing value and neglected imbalanced class problems within the dataset. The dataset has more than 5% missing values and has a significant amount of imbalance. The class that has the most amount of data has around 64% of the data, but the least amount of data only has 0.6% of the data. On the other hand, the dataset’s missing values are randomly scattered among the data. This research proposed a preprocessing process to combine Multiple Imputation by Chained Equations (MICE) and Synthetic Minority Oversampling Technique (SMOTE) and tested it with three machine learning methods such as Random Forest, Support Vector Machine, and K-Nearest Neighbour. As a result, the g-mean metrics for those three machine learning approaches improved from 0.2 percent to 3.8 percent.