An Analysis of Several Machine Learning Algorithms for Imbalanced Classes
Soma Datta, Anuprabha Arputharaj · 2018
Imbalanced data typically refers to classification problems when classes are not represented equally. In real world applications, most classification datasets do not have exactly equal number of instances and the problem arises where the class is imbalanced. Most Machine learning algorithm works best when the number of instances of each class are roughly equal but there are only specific algorithms to deal with the imbalanced classes. This survey research is mainly used to assess the thoughts and opinions of several authors regarding the imbalanced class in data mining. This survey focuses on information from multiple datasets and it aims to obtain several perspectives about imbalanced class. The survey is implemented by analysing detailed reports on several datasets from the UCI-Machine Learning Repository which is the centre for Machine Learning and Intelligent Systems. This study includes 20 datasets and their methodologies from 26 articles to do a comparative study on imbalanced class problems from fuzzy classification, decision trees, association mining, and ensemble methods. Class imbalance problem is extremely common in practice and is observed in various disciplines including medical diagnosis, fraud detection, anomaly detection, oil spillage detection, facial recognition, etc. However, this problem affects machine learning due to having disproportionate number of class instances in practice and due to its prevalence, several approaches are studied to deal with this problem. This study aims to exhibit one such approach for handling different datasets.