Enhancing Fairness and Performance: Scalable Hybrid Solutions for Class Imbalance in Big Data Analytics

B H Puneetha, Manoj Kumar M V, Prashanth B S, Ajay Kumara M.A. · Procedia Computer Science · 2025

Recent advancements in handling class imbalance include hybrid approaches that combine data-level and algorithmic-level techniques. However, challenges such as computational efficiency and adaptability to dynamic datasets remain to be prominently addredssed. This research paper introduces novel hybrid techniques that combine both data-level and algorithmic-level approaches to address class imbalance in large datasets. These techniques are designed to improve class distribution and processing efficiency while enhancing prediction accuracy, particularly for minority classes. The proposed model applies K-means SMOTE for over-sampling the minority class and centroid-based undersampling for the majority class, followed by Classification using XGBoost. Evaluations on multiple datasets demonstrate significant improvements in precision, recall, AUC-ROC, and F1-score compared to traditional methods. In addition to enhancing robustness and fairness, the approach effectively addresses challenges related to scalability, data diversity, and dynamic data sources in big data analytics. The hybrid technique achieved a 10-20% improvement in key performance metrics, including precision, recall, F1-score, and AUC-ROC, across various UCI datasets such as Diabetes, Hepatitis, and German Credit, confirming the model’s superiority in handling class imbalance and improving overall Classification performance.

Read the paper · More papers on PaperTik