Accuracy Analysis of Supervised and Unsupervised Techniques on Breast Cancer Datasets
Elham Jahanpeikar, Mahdi Ebrahimi · 2022
Modern medicine benefits extensively from technology as a valuable assistant to diagnose diseases at early stages and cure them with less complication and more success. In the health sector, cancer diseases undeniably, possess the most variety and complexity of all diseases. Therefore, this paper focuses on breast cancer and applies supervised, unsupervised, and deep learning as one candidate from each major machine learning field to evaluate different models and analyze their outcomes. We aim to answer the following questions: Which dataset produces the best result and improves its metrics and performance using machine learning optimization techniques? Will there be any definite decision on which machine learning models or algorithms have superiority over the others? The original and diagnostic datasets are trained by Logistic Regression, K-means clustering, and multi-layer perceptron or artificial neural network model (ANN). Then, unsupervised techniques, heatmaps, and Principal Component Analysis (PCA) are used to reduce dimensionality and concise the dataset for any probable improvements. The original dataset produced better results for the machine learning models, and ANN obtained the best accuracy score. The comprehensive and systematic calculation of the metrics and indexes of the breast cancer datasets and the thorough optimization by unsupervised technics is the novelty of this research. The comparison between these two datasets has not been approached before. The clustering by K-means creates novel visualization of the datasets, which could give the experts in the field ideas of the cancerous mass's characteristics.