Abstract C028: Next-generation breast cancer detection: Harnessing AI, machine learning, and deep learning neural network model for highest accuracy, non-invasive, cost-effective early intervention and superior survival outcomes
Gokul P. Kesav · Cancer Epidemiology Biomarkers & Prevention · 2025
Abstract Abstract: Breast cancer is the most common cancer affecting 2M women worldwide, and second leading cause of death killing 0.5M women. Early detection can save 4x costs and increase the survival rate by 99%. Most biopsies are intrusive, painful and expensive to patients as they are done via surgery or core needle. My experiment is to prove that breast cancer can be detected with 100% precision using the Machine Learning (ML) model I developed, and with data collected from the least intrusive, least painful, least expensive biopsy technique called Fine Needle Aspiration (FNA). To accurately diagnose breast masses from FNA, Dr. Wolberg (Oncology Surgeon) University of Wisconsin published a dataset of 569 patients. Using this anonymized patient dataset, I developed various Deep Learning AI/ML Models and tweaked the hyper parameters to develop the best model to detect cancer. My model in Python based on logistic regression has the highest precision to predict breast cancer at 100%, and far exceeded any other models based on the same data available online. Compared to other models published in Kaggle, my model has test scores consistently with the highest precision (100%),highest recall (90.4%), highest accuracy (96.49%) using the least invasive biopsy technique (FNA). Experimental Methods: The procedure for my experiment is as follows: 1) Collect dataset for breast cancer tumors from a reliable resource. In this case, I got the dataset from Dr. Wolberg research in breast cancer. 2) Analyze and understand the data - this step involves feature engineering. Understand how the different features are correlated with the output and also between each other. This helps improve model performance, interpretability, generalization, and computational efficiency. 3) Data preprocessing and clean up of data to present only relevant information. This involves removing the features that may cause the model to overfit. 4) Train/Develop a Machine Learning(ML) model to classify the data using the different features. I explored the logistic regression models available on AWS SageMaker 5) Validate the Machine Learning model to ensure it is generalized for the data and is not overfitting. I used ADAM optimizer (Adaptive Moment Estimation) to efficiently update the weights, biases of my neural network model during training by dynamically adjusting the learning rate for each parameter based on its gradient history. 6) Test the AI/ML Deep Learning model, report the results 7) Keep changing the hyperparameters (like epochs), model weights and biases to get different models and conclude the experiment when the model testing and validation loss values are lowest (anything beyond this will cause the model to overfit). Results and Conclusions: Accuracy: 96.49% Precision: TP/(TP+FP) : 100% Recall: TP/(TP+FN) : 90.4% ROC-AUC (Receiver Operating Characteristic - Area Under the Curve): 0.98 My novel method of predicting breast cancer under highest precision (100%), has the highest recall and accuracy than other models available publicly. Citation Format: Gokul P. Kesav. Next-generation breast cancer detection: Harnessing AI, machine learning, and deep learning neural network model for highest accuracy, non-invasive, cost-effective early intervention and superior survival outcomes [abstract]. In: Proceedings of the 18th AACR Conference on the Science of Cancer Health Disparities; 2025 Sep 18-21; Baltimore, MD. Philadelphia (PA): AACR; Cancer Epidemiol Biomarkers Prev 2025;34(9 Suppl):Abstract nr C028.