EasyEnsemble and Feature Selection for Imbalance Data Sets
Tianyu Liu · 2009
There are many labeled data sets which have an unbalanced representation among the classes in them. When the imbalance is large, classification accuracy on the smaller class tends to be lower. In particular, when a class is of great interest but occurs relatively rarely such as cases of fraud, instances of disease, and so on, it is important to accurately identify it. Here we propose a novel algorithm named MIEE (mutual information based feature selection for EasyEnsemble) to treat this problem and improve generalization performance of the EasyEnsemble classifier. Experimental results on the UCI data sets show that MIEE obtain better performance, compared with the asymmetric bagging and EasyEnsemble.