Latin Hypercube Sampling Approach to Improve K-Nearest Neighbors Performance on Imbalanced Data

Khairul Umam Syaliman, Adli Abdillah Nababan, Miftahul Jannah, Arif Hamied Nababan, Ryan Dhika Priyatna, Erwin Panggabean · 2023

Imbalanced class is a common issue encountered in real-world datasets. Oversampling is a technique used to tackle imbalanced classes, with the Synthetic Minority Oversampling Technique (SMOTE) being the most popular method. However, SMOTE has limitations, including diversity problems, inconsistent synthesized data, and the generation of outliers or noise. To address these challenges, one proposed solution is to apply oversampling using the Latin Hypercube Sampling (LHS) approach. The LHS approach involves selecting the majority class based on the imbalanced ratio, synthesizing data using LHS, and then selecting the synthesized data. The wine quality and glass identification datasets were used in this study. The LHS-based data synthesis has been proven to enhance the performance of the k-Nearest Neighbors (k-NN) algorithm. The experimental results indicate an average improvement in performance metrics such as accuracy, precision, recall, and F1-Score, with values of 0.1656, 0.35805, 0.3728, and 0.36885, respectively.

Read the paper · More papers on PaperTik