Multiple feature construction for effective biomarker identification and classification using genetic programming
Soha Ahmed, Mengjie Zhang, Lifeng Peng, Bing Xue · 2014
Biomarker identification, i.e., detecting the features that indicate differences between two or more classes, is an important task in omics sciences. Mass spectrometry (MS) provide a high throughput analysis of proteomic and metabolomic data. The number of features of the MS data sets far exceeds the number of samples, making biomarker identification extremely difficult. Feature construction can provide a means for solving this problem by transforming the original features to a smaller number of high-level features. This paper investigates the construction of multiple features using genetic programming (GP) for biomarker identification and classification of mass spectrometry data. In this paper, multiple features are constructed using GP by adopting an embedded approach in which Fisher criterion and p-values are used to measure the discriminating information between the classes. This produces nonlinear high-level features from the low-level features for both binary and multi-class mass spectrometry data sets. Meanwhile, seven different classifiers are used to test the effectiveness of the constructed features. The proposed GP method is tested on eight different mass spectrometry data sets. The results show that the high-level features constructed by the GP method are effective in improving the classification performance in most cases over the original set of features and the low-level selected features. In addition, the new method shows superior performance in terms of biomarker detection rate.