Data Mining as a Knowledge Extraction Tool and an Application on Decision Tree-Based Algorithms
Fatma Önay Koçoğlu · 2022
In today’s information society, especially the knowledge has come to the forefront; changes have occurred in the habits of individuals and societies, in the business models and processes of institutions and organizations. Developing database and data warehouse technologies allows the collection, storage, and management of more and complex data with less cost. This recorded data becomes a great resource in the effort to access knowledge. This huge recorded data must be processed and analyzed with appropriate techniques and converted into knowledge quickly and effectively. Statistical and mathematical based machine learning and data mining methods that enable data analysis have been developed. Data mining is the job of extracting knowledge from the data stack. There are many situations that affect data mining model performance. One of the situations is the imbalanced distribution of the class attribute field values in the data set. Within the scope of this study, it is aimed to observe the performance change of various decision tree-based algorithms used in the data mining process against data with balanced and imbalanced class distribution and to show what the implications can be obtained with a data mining study for enterprises. C4.5, C5.0, C5.0 (boosted), Gini, and Random Forest algorithms have been used to develop the classification model and undersampling to balance the data set. The results show that it is important to eliminate the unbalanced class distribution in order to evaluate the model in real terms.