Class Imbalance of Bio-Medical Data by Using PCA-Near Miss for Classification

Praveen Tumuluru, Ravuri Daniel, Gundabathula Mahesh, Kalavala Deekshitha Lakshmi, Pakalapati Mahidhar, M. V. Manoj Kumar · 2023

Class imbalance is a significant problem in many real-time domains dealing with healthcare data, which leads to poor classification performance and may have severe consequences. This work addresses the class imbalance problem in healthcare data, which often leads to poor classification performance and potentially serious consequences. This study proposes a hybrid PNM model that integrates the Near-Miss sampling technique and Principal Component Analysis to overcome this challenge. This study aims to investigate the effectiveness of the PNM model in improving classifier performance and compare it with several baseline classifiers. Further, the experiments are conducted on a real-world healthcare dataset containing various healthcare records. The results showed that the PNM model outperformed the baseline classifiers regarding the precision, recall, F1 score, and area under the curve (AUC). Specifically, the PNM model achieved a precision of - 0.85, recall of - 0.82, F1 score - 0.83, and AUC - 0.91, while the best-performing baseline classifier achieved a precision of 0.72, recall of 0.64, F1 score of 0.68, and AUC of 0.83. Our study demonstrates that the PNM model offers a promising approach to addressing the class imbalance in healthcare data and improving classifier performance. Integrating the Near-Miss sampling technique and Principal Component Analysis enables the model to achieve a better balance among the majority and minority classes, resulting in more accurate classification. The PNM model has the potential to be applied to various healthcare domains, such as disease diagnosis, patient risk stratification, and treatment prediction.

Read the paper · More papers on PaperTik