Boosting Classifier Accuracy on Imbalanced Data with an Enhanced SMOTE Approach

G. Sajiv, Natarajan Meenakshisundaram · 2025

Imbalanced datasets present considerable hurdles in machine learning, especially in fields such as healthcare, where precise identification of minority class instances is essential. Conventional resampling techniques, such as SMOTE (Synthetic Minority Oversampling Technique), frequently neglect feature-specific significance, resulting in inferior classifier efficacy. This paper introduces Feature-Weighted SMOTE (FW-SMOTE), an innovative adaptation of SMOTE that use feature importance to direct synthetic data generation, hence producing more significant and representative samples. We assess the efficacy of FW-SMOTE in comparison to regular SMOTE utilizing three prevalent classifiers: Random Forest (RF), Support Vector Machine (SVM), and Logistic Regression (LR), on a dataset concerning cervical cancer risk factors. Performance is evaluated by precision, recall, and F1-score, prior to and after to the use of each resampling strategy. Results indicate that FW-SMOTE regularly enhances classifier performance, especially for SVM and LR, by producing synthetic samples that more accurately represent the data distribution. Furthermore, comparative bar graphs and tabular results underscore the efficacy of FW-SMOTE in addressing imbalanced datasets, indicating its capacity to enhance classification jobs in essential fields. Our findings indicate that FW-SMOTE is a resilient and efficient method for improving model performance, facilitating more precise predictions in practical applications.

Read the paper · More papers on PaperTik