Rough–Granular Approach in Imbalanced Bankruptcy Data Analysis
Katarzyna Borowska, Jarosław Stepaniuk · Procedia Computer Science · 2022
During the global economic crisis data were collected on the financial condition of Polish enterprises. Correct analysis of this type of data gives a chance to draw conclusions and avoid the problem of bankruptcy in the future. Unfortunately, the collected data have imbalanced characteristics. Hence, they require a dedicated processing strategy. In this paper, we propose to use an optimized version of our rough–granular approach (RGA) to improve the classification efficiency of bankruptcy data. The RGA oversampling algorithm generates new minority class samples in specific areas of the feature space taking into consideration additional difficulty factors. Moreover, the method in the filtering step provides removal of generated ambiguities. This approach has been successful with bankruptcy data. The study showed in most cases an improvement in the detection rate of the minority class compared to the classification without preprocessing and the SMOTE technique. The RGA algorithm was confirmed to be effective in terms of AUC and recall measures on imbalanced data with complex distribution. The accuracy metric proved to be unreliable for data with these characteristics. XGBoost was used as classifier. In our experiments, we focused on an in-depth analysis of the types of information granules present in the dataset and their impact on classification performance.