A New approach for Classification of Highly Imbalanced Datasets using Evolutionary Algorithms
Satyam Maheshwari, Jitendra Agrawal, Sanjeev Sharma · 2011
Today's most of the research interest is in the application of evolutionary algorithms. One of the examples is clas- sification rules in imbalanced domains. The problem of Imbalanced data sets plays a major challenge in data mining community. In imbalanced data sets, the number of instances of one class is much higher than the others, and the class of fewer represent- atives is of more interest from the point of the learning task. Traditional Machine Learning algorithms work well with balanced data sets, but not able to deal with classification of imbalanced data sets. In the present paper we use different operators of Ge- netic Algorithms (GA) for over-sampling to enlarge the ratio of positive samples, and then apply clustering to the over-sampled training dataset as a data cleaning method for both classes, removing the redundant or noisy samples. The proposed approach was experimentally analyzed and the experimental results shows an improvement in the classification measured as the area under the receiver operating characteristics (ROC) curve. —————————— a —————————— 1 I HE problem of imbalanced data-sets occurs when the majority class has a large percent of the samples, while minority class occupies a small part of all samples. Such a condition pose challenges for classical machine learning algorithms that are designed to optimize oTverall classifica- tion accuracy. Imbalanced datasets exists in many domains such as medical applications (1), risk management (2), face recognition (3) and information technology, and so on. In these domains, minority class is of more interest than majori- ty class. In imbalanced data sets, the traditional way of max- imizing overall performance will often fail to learn anything useful about the minority class, because of the dominating effect of the majority class. A learner can probably achieve 99% accuracy with ease, but still fail to correctly classify any rare examples. Therefore, analyzing the imbalanced data sets (IDS) problem requires new and more adaptive methods than those used in the past. In this paper we over-samples the minority class by mutation and crossover operators to decrease the imbal- ance ratio and then using clustering for both classes to delete redundant and noisy samples. Thus, by combining the both method the samples of interest are remained, improving the computational efficiency. The contribution is organized as follows: Section 2 in- troduces the problem of imbalanced data sets, describing