Enhanced SMOTE Strategy for Handling Imbalanced Data in Machine Learning Classification
Shivam Tiwari, Satvik Vats, Barkha Bhardwaj, Pramod Kumar Vishwakarma, Yogesh Singh Rathore · 2023
Most real-time datasets have class instances scattered across the data in a non-uniform fashion. Certain classes have an abnormally high number of instances when compared to the rest of the classes. The term "class imbalance problem" is used to describe this occurrence (CIP). Due to its inaccuracy in projecting weak class data samples, this skewed information impacts the reliability of forecasts as a whole. Many businesses make use of CIP's data mining experts who employ CIP. Both ML and deep learning have substantial difficulties with the classification of skewed data (DL). Several researchers are curious in how to overcome the major hurdle presented by the widespread adoption of sample techniques to improve classifier performance. In this research, we provide a survey of state-of-the-art approaches currently being used in studies of data classification, and we contrast and compare the numerous algorithms used at each stage of the process. Talk about the challenges and limitations that come up in investigations of imbalanced data classification. Taken as a whole, this work not only provides an in-depth analysis of the uneven data categorization industry but also inspires more study in this space.