Classification of Imbalanced Datasets Using Various Techniques along with Variants of SMOTE Oversampling and ANN
M. Shrinidhi, T.K. Kaushik Jegannathan, R. Jeya · Advances in science and technology · 2023
Using Machine Learning and / or Deep Learning for early detection of diseases can help save people’s lives. AI has already been making progress in healthcare as there are newer and improved software to maintain patient records, produce better imaging for error free diagnosis and treatment. One drawback working with real-life datasets is that they are predominantly imbalanced in nature. Most ML and DL algorithms are defined keeping in mind that the dataset is equally distributed. Working on such imbalanced datasets cause the models to end up having high type-1 and type-2 error which is not ideal in the medical field as it can misdiagnose and be fatal. Handling class imbalance thus becomes a necessity lest the ML/DL model fails to learn and starts memorizing the features and noises belonging to the majority class. PIMA Dataset is one such dataset with imbalances in classes as it contains 500 instances of one type and 268 instances of another type. Similarly, the Wisconsin Breast Cancer (Original) Dataset is also a dataset containing imbalanced data related to breast cancer with a total of 699 instances where 458 instances are of one class (Benign tumor images) while 241 instances belong to the other class (Malignant tumor images). Prediction/detection of onset of diabetes or breast cancer with these datasets would be grossly erroneous and hence the need for handling class imbalance increases. We aim at handling the class imbalance problem in this study using various techniques available like weighted class approach, SMOTE (and its variants) with a simple Artificial Neural Network model as the classifier.