Selecting the Appropriate Ensemble Learning Approach for Balanced Bioinformatics Data

David J. Dittman, Taghi M. Khoshgoftaar, Amri Napolitano · The Florida AI Research Society · 2015

Ensemble learning (process of combining multiple models into a single decision) is an effective tool for improving the classification performance of inductive models. While ideal for domains like bioinformatics with many challenging datasets, many ensemble methods, such as Bagging and Boosting, do not take into account the high-dimensionality (large number of features per instance) that is commonly found in bioinformatics datasets. This work seeks to observe the effects of two relatively new ensemble learning methods (Select-Bagging and Select-Boosting: the Bagging and Boosting approaches with feature selection implemented within each iteration of their algorithms) on a series of seven balanced (greater than a 43.50% minority class distribution) bioinformatics datasets. Additionally, we included the results when no ensemble approach is implemented (denoted as No-Ensemble) so that we can observe the full effects of ensemble learning. In order to test the three approaches we use three feature rankers, four feature subset sizes, and two classifiers. The results show that Select-Bagging is the top performing ensemble approach and statistical analysis confirms that Select-Bagging is significantly better than No-Ensemble and better (though not significantly) than Select-Boosting. Our recommendation is that SelectBagging is an excellent choice for improving classification performance for bioinformatics datasets. To our knowledge, this work is the first empirical study focused exclusively on balanced bioinformatics datasets that investigated the effects of ensemble learning and utilizes Select-Bagging and introduces Select-Boosting.

Read the paper · More papers on PaperTik