Feature Selection based Performance Comparison of Classifier Models for an Imbalanced Dataset: Early Diagnosis of Symptoms for Ovarian Cancer
Suhas Suresh Ambekar, Natashaa Kaul, Manoj Hudnurkar, Niharika Niharika · 2023
This research aims to reveal the evidence that can be used to narrow the knowledge gap, understand the early signs, and undergo prediagnosis for Ovarian Cancer patients. After the exploratory data analysis, classifier models are applied to the ovarian cancer dataset. The performance of all the classifiers before and after dimensionality reduction was analyzed based on confusion matrix parameters. Applying scoring methods like Information Gain and Gini Decrease for Feature Selection, Principal Component Analysis, and an integrated approach is used for feature extraction. A comparative analysis of classifiers was done, and the performance results were evaluated using the weighted TPR-TNR method as the dataset is imbalanced. The scoring feature selection method has been used to rank the features, which aims to reduce the data's dimensionality to eight for larger datasets, allowing features with maximum information and providing the best attributes by using this methodology, an accuracy of 99.8% on a different validation dataset could be achieved. Gradient Boosting, ANN, Adaboost, and Random Forest are the most recommended classifier algorithms. This study indicates that out of nine classifier models, two models, ANN and Gradient Boosting, significantly improve outcomes with fewer dimensions. This will also help the healthcare industry regulate unnecessary treatment delays by focusing on a limited number of features.