Accuracy Enhancement of Machine Learning Model by Handling Imbalance Data
Manimuthu Ayyannan · 2024
To enhance the performance of a Machine Learning (ML) model on class-imbalanced data, this study will analyze different approaches toward class balancing. High-class imbalance is inherently prevalent in numerous practical applications such as fraud detection and rare disease diagnosis, making successful categorization with imbalanced data a significant field of study. In addition, the problem is exacerbated by severely imbalanced data because most ML models will show bias towards the dominant class and, in the worst cases, neglect the minority class completely. There has been relatively limited research in the domain of class balancing methods, despite recent breakthroughs in handling imbalance data approaches and growing popularity. In this research, we create and evaluate several methods for balancing data, including Random Under Sampling (RUS), Random Over Sampling (ROS), Adaptive Synthetic Sampling (ADASYN), Synthetic Minority Oversampling Technique (SMOTE), and Synthetic Minority Oversampling Technique Tomek (SMOTETomek). We use credit card fraud detection data for our research. This data went through a process of data balancing treatment. Then, the Support Vector Machine (SVM) ML approach is trained with the help of evenly distributed data. SVM mode results for all of the aforementioned balancing data types are evaluated, and the optimum method for handling imbalanced data is determined. The outcome demonstrates that the SMOTETomek method is superior to others in terms of accuracy.