Imbalanced Dataset Problem in Classification Algorithms
Nur Çürükoğlu · 2019 1st International Informatics and Software Engineering Conference (UBMYK) · 2019
Nowadays, companies continuous calculations and research with the available data to minimize the cost of personnel and time. Within the company, they provide an environment in which employees can enter their suggestions for improvement or complaints with the purpose of provide better service. Accordingly, use of "Personal suggestion systems" has been increasing by corporate companies during the recent years. In order to automate suggestion systems and make some analyzes, commonly used and developing machine learning technologies are used today. On the other hand, one of the problems encountered of machine learning as is the data set problem having imbalanced distribution. In the real world, data sets have imbalanced are quite a much. In this study, the results of ROS, RUS, SMOTE and ADASYN methods on the classification algorithms were analyzed and the best method "SMOTE" and "Gradient Boosting Classifier" classification algorithm which gave the best results were preferred.