On Producing Balanced Fuzzy Decision Tree Classifiers
Keeley A. Crockett, Zuhair A. Bandar, James D. O’Shea · 2006
This paper investigates a new approach to creating robust fuzzy classifiers that are impartial to the imbalance problem within data sets. The approach uses raw real-world data without the need for sampling or the creation of synthetic examples. The aim is to achieve common currency between the actual classification accuracy and the distribution of this accuracy between the outcome classes. The proposed method first uses a fuzzy inference algorithm (FIA) to construct a fuzzy classifier from a crisp C4.5. A genetic algorithm (GA) is then used to optimize the degree of fuzziness within the classifier. The GA's fitness function consists of two components: classification accuracy and the distribution (or balance) of this accuracy between the outcomes. Both components are optimised concurrently. Four alternative fitness functions are defined, each of which applies different penalties on the classification accuracy depending on a weighting associated with the balance component. The method is then applied to three real world data sets. The results show that it is possible to attain a fuzzy classifier which exhibits both good performance and balance between outcomes regardless of any imbalance within the data set.