A Clustering Based Priority Driven Sampling Technique for Imbalance Data Classification

Iftakhar Ali Khandokar, Abdullah All Tanvir, Tanvina Khondokar, Nabila Tabassum Jhilik, Swakkhar Shatabda · 2022

Classification of Imbalance data is one of t he most vital tasks in the field of machine learning because most of the real-life datasets available have an imbalanced distribution of class labels. The effect of imbalanced data is severe where the predictive model trained on the imbalanced data faces some unprecedented problems like overfitting where t he model gets biased towards the majority target class. Many techniques have been proposed over time to deal with the imbalanced distribution caused by problems like oversampling and undersampling where oversampling isn't able to match the performance acquired by the undersampling method. One such baseline method is clustering the majority of data into multiple clusters and then randomly sampling some of the redundant data but we believe that randomly sampling the data sample might open the loophole to losing informative data samples. So, in this work, we would like to propose two clustering-based priority sampling methods which manage to boost the performance of the predictive model compared to the clustering-based random sampling techniques.

Read the paper · More papers on PaperTik