An ANN-Based Resampling Approach for Handling Imbalance and Overlapped Data
Talha Ramzan, Noshina Tariq, Umar Farooq, Husin Jazri, Humaira Ashraf, Jan Sher Khan, Pranav Srinivas Kumar · 2024
Imbalanced datasets in class are common across numerous zones, including banking, security, fraud detection, and health. Most supervised learning methods are biased towards the dominant class when dealing with imbalanced datasets. The learning effort increases when there is a mix of examples from various classes. As a result, by deleting possibly overlapping data points, a resampling structure for controlling class imbalance in binary datasets has been developed. The methods are intended to find and remove majority class instances from the overlapping area. At the data level, solutions for unbalanced class situations try to change the class distribution. Resampling data via under-sampling or oversampling, which decreases the majority class appearances and raises the minority class instances, is a popular procedure. The challenge may be solved at the algorithmic level by developing new learning algorithms or altering current ones. The benefit of algorithm-level solutions is that they directly include the user's choices in the model. Accurate identification and eradication of these instances maximize the visibility of the minority class instances while minimizing data removal, resulting in less information loss. Techniques based on Artificial Neural Networks (ANN) reach good mean average precisions. The study aims to enhance the resampling approach for decomposition using ANN with undersampling and oversampling techniques like (Nearest Neighbour Random UnderSampling (NN-RUS), Near Miss UnderSampling (NMUS), Cluster Centroid Undersampling (CCUS) and SMOTE OS) that will become visible for minority class instances. Extensive tests were conducted utilizing both simulated and real-world datasets. The results reveal that state-of - the-art approaches perform similarly across various popular criteria, with extraordinary and statistically significant gains in sensitivity and accuracy of the classifier without compromising the minority class instances of a dataset.