ROAFS: Interpretable classification of imbalanced medical data based on random oversampling and AFS decision trees
XuLi Tan, Xun Gong, Siyu Qin, Xinxin Li, Wenjuan Jia · Mathematical Foundations of Computing · 2025
Imbalanced data presents a critical challenge in the medical domain, significantly influencing the accuracy of disease diagnosis and treatment. Traditional machine learning algorithms often favor the majority class. Common approaches to mitigate this bias include augmenting minority class samples, reducing majority class samples, or leveraging ensemble classifiers. However, in medical applications, these methods often compromise data authenticity and information density, or require extensive computational resources and intricate parameter tuning to achieve optimal results. To address these challenges, this paper introduces a method named ROAFS that integrates random over-sampling techniques with axiomatic fuzzy set theory based decision trees. This approach effectively improves both classifier performance and interpretability. Random over-sampling is utilized to balance the imbalanced datasets, which are combined with fuzzy decision trees to generate clear semantic descriptions for each class. Experiments conducted on 10 datasets with 5-fold cross-validation show that ROAFS maintains high classification accuracy. It also provides intuitive and meaningful semantic explanations, outperforming existing methods. This approach offers a practical and interpretable solution to the problem of imbalanced data, showing significant potential for real-world medical applications.