Optimisation of Cancer Status Prediction Pipelines using Bio-Inspired Computing
Mariel Barbachan e Silva, Pedro Henrique Narloch, Márcio Dorn, Pilib Ó Broin · 2021
Cancer is one of the leading causes of death globally, and early detection is a fundamental factor in improving patient outcomes. The advent of high-throughput genetic profiling techniques in the last few decades has led to an explosion of genetic data related to cancer. Machine learning methods, and classification algorithms in particular, have been used to find underlying patterns in cancer data and make diagnostic predictions. The addition of feature selection to classification pipelines can lead to improvements in predictive capabilities, since the removal of non-important features benefits the construction of classification models. We developed a classification pipeline for cancer status prediction composed of a feature selection step with SelectKBest and an ensemble classifier system with five popular supervised learning algorithms. We used three bio-inspired optimization techniques to select the optimal sets of hyperparameters for the classification pipeline and compared these approaches on three cancer microarray datasets. The results indicate that the optimized pipelines have better predictive performance in all but one of the experiments compared to the ensemble alone.