Imbalanced Classification Using Genetically Optimized Random Forests

Todd Perry, Mohamed Bahy Bader-El-Den · 2015

Class imbalance is a problem that commonly affects 'real world' classification datasets, and has been shown to hinder the performance of classifiers. A dataset suffers from class imbalance when the number of instances belonging to one class outnumbers the number of instance belonging to another class. Two ways of dealing with class imbalance are modifying the dataset to reduce the number of instances belonging to the majority class(es) (known as resampling), or allowing the classifier to penalize misclassifying the minority class(es) more than the majority class(es), this can be done by implementing a cost matrix. This paper attempts to improve the classification performance of the Random Forest classifier on imbalanced datasets by exploiting these two techniques, to do this a genetic algorithm is employed to find optimal parameters. Results are compared to commonly used classification algorithms.

Read the paper · More papers on PaperTik