Cluster Based Under-Sampling for Unbalanced Cardiovascular Data
M. Mostafizur Rahman, Darryl N. Davis · 2013
Abstract—Most medical datasets are not balanced in their class labels. Indeed in some cases it has been no ticed that the given class labels do not accurately represent characteristics of the data record. Most existing classification methods tend not to perform well on minority class examples when the dataset is extremely imbalanced. This is because they aim to optimize the overall accuracy without considering the relative distribution of each class. In this paper we propose a cluster based under sampling technique that solves the class imbalance problem for our cardiovascular data. It shows significant better performance than existing methods.