A Class-Imbalanced Study with Feature Extraction via PCA and Convolutional Autoencoder

Zahra Salekshahrezaee, Joffrey L. Leevy, Taghi M. Khoshgoftaar · 2022

It is inherently challenging to train a machine learning algorithm on a class-imbalanced dataset. Under conditions of high dimensionality, this training process can become even more difficult due to the large number of features in the dataset. During preprocessing, data sampling is commonly used to address class imbalance and feature extraction is frequently used to reduce the number of dataset features. In this study, we explore the use of these two preprocessing activities before passing on the data to four ensemble classifiers (Random Forest, CatBoost, LightGBM, and XGBoost). With reference to feature extraction, the Principal Component Analysis (PCA) and Convolutional Autoencoder (CAE) methods are evaluated. With regard to data sampling, the Random Undersampling (RUS) and Synthetic Minority Oversampling Technique (SMOTE) methods are evaluated. Classification performance is measured with the Area Under the Receiver Operating Characteristic Curve (AUC) metric. Our results indicate that the implementation of the RUS method followed by the CAE method leads to the best classification performance.

Read the paper · More papers on PaperTik