Evaluating the Effectiveness of Smote for Imbalanced Data Expansion and Its Impact on Classification Accuracy
Shaik Mohammed Imran, Angelina Geetha · 2024
Big data classification technology has emerged together with artificial intelligence, offering valuable support for studies on auxiliary diagnostics in medicine. Medical big data is frequently unbalanced because of the various conditions in the many sample collections. Many popular learning methods have their classification performance hindered by the class-imbalance problem. While the SMOTE algorithm's random sample point generation feature could lead to an increase in the imbalance rate, the blindness of parameter selection and marginalization formation make its deployment problematic. This research seeks to remedy the situation by presenting a normal distribution-based SMOTE algorithm that is superior to its predecessor. The extra sample points will be distributed more fairly, which will prevent larger data portions from being underrepresented. Experiments show that when applied to imbalanced datasets like Pima, WDBC, WPBC, Ionosphere, and Breast Cancer, the new method outperforms the original SMOTE algorithm in terms of classification performance. This was demonstrated in Wisconsin. Our thorough testing also revealed that maintaining the distribution properties of the original data through the selection of appropriate parameters in the suggested method results in the best classification impact.