Tumor Detection of Breast Tissue Using Random Forest with Principal Component Analysis

Gul Muhammad Soomro, Said Krayem, Zaira Hassan Amur, Bronislav Chramcov, Roman Jašek, Ismail Noordin · 2023

When it comes to cancer-related mortality, breast cancer is the most common and predominant kind in women; lung cancer is the most common. It ranks second overall. Global scientists have been working very hard to fight this illness for a long time. Furthermore, significant advancements have been achieved in the fields of machine learning and data mining for the extraction and synthesis of insightful knowledge, even from extremely complicated data sources. Machine learning models may carry out a number of functions, including clustering, classification, and prediction, by using the knowledge obtained from data. In this study, we examine the relationship between a dataset's several features and a diagnosis of breast cancer. We use five different variables extracted from X-ray images in the dataset to predict the existence of breast cancer using a supervised learning classification technique called Random Forest. The open-source web repository Kaggle served as the source of this dataset. We use Principal Component Analysis (PCA), a dimensionality reduction method, on the data prior to putting the Machine Learning algorithm into practice. We then assess the performance of the Machine Learning algorithm using measures like accuracy, precision, recall, F1-score, and support. Additionally, we compare the model's output with Random Forest and examine the performance metrics it produces with and without PCA analysis.

Read the paper · More papers on PaperTik