Machine Learning Techniques Improving the Box–Cox Transformation in Breast Cancer Prediction
Sultan S. Alshamrani · Electronics · 2025
Breast cancer remains a major global health problem, characterized by high incidence and mortality rates. Developing accurate prediction models is essential to improving early detection and treatment outcomes. Machine learning (ML) has become a valuable resource in breast cancer prediction; however, the complexities inherent in medical data, including biases and imbalances, can hinder the effectiveness of these models. This paper explores combining the Box–Cox transformation with ML models to normalize data distributions and stabilize variance, thereby enhancing prediction accuracy. Two datasets were analyzed: a synthetic gamma-distributed dataset that simulates skewed real-world data and the Surveillance, Epidemiology, and End Results (SEER) breast cancer dataset, which displays imbalanced real-world data. Four distinct experimental scenarios were conducted on the ML models with a synthetic dataset, the SEER dataset with the Box–Cox transformation, a SEER dataset with the logarithmic transformation, and with Synthetic Minority Over-sampling Technique (SMOTE) augmentation to evaluate the impact of the Box–Cox transformation through different lambda values. The results show that the Box–Cox transformation significantly improves the performance of Artificial Intelligence (AI) models, particularly the stacking model, achieving the highest accuracy with 94.53% and 94.74% of the F1 score. This study demonstrates the importance of feature transformation in healthcare analytics, offering a scalable framework for improving breast cancer prediction and potentially applicable to other medical datasets with similar challenges.