Enhancing Breast Cancer Diagnosis: A Random Forest-Based Approach to Machine Learning

Priyanka V. Deshmukh, Aniket Kailas Shahade · 2025

Breast cancer is most lethal cancers in women globally making it important to design effective and precise diagnostic tools. The old ways of diagnosing breast cancer include mammography as well as biopsies and may be expensive or have low accuracy. To stimulate the utilization of machine learning (ML) as an alternative approach to improve the chance of detection as early as possible and classification of breast cancer, this outlines the following. The diagnostic model in this research work uses Wisconsin Diagnostic Breast Cancer to train and test the machine learning solution. Analysis revealed important features and critical features’ correlation in the dataset, which was done as an exploratory data analysis (EDA). The dataset includes 569 instances and 30 continuous attributes based on cell nucleus measurements plus one class label that identifies the tumor type, malignant or benign. Before developing Machine Learning models all sorts of pre-processing were done such as missing values are handled, categorical variables were encoded and features were standardized. Random Forest Classifier was used because of its stability, features and great predictive accuracy. Using the model, an accuracy of 96.5% was obtained with precision or recall of 95.7% and recall of 97.3%. Among these, the highest recall score is a critical value that shows the high ability of the designed model to detect malignant cases and minimise false negative results. Thereby, the literature results show that the Random Forest Classifier has the potential to become an effective tool for diagnosing breast cancer as an easily accessible, rapid, and accurate diagnostic method. Further development of these results in future studies could involve various works in the field of ensemble learning and clinical data for higher accuracy and clinical relevance.

Read the paper · More papers on PaperTik