Gene selection and decision tree based classification for cancerous sample detection
Sunanda Das, Asit Kumar Das · International Journal of Biomedical Engineering and Technology · 2016
Generally, gene expression data are of high-dimensional which cause degradation of the performance of gene data analysis for disease prediction. Therefore, it is a big issue for the traditional classifiers to perform well on high-dimensional microarray data where the number of genes far exceeds the number of samples. In the proposed work, initially, Pearson's correlation coefficient is computed between every pair of genes and based on these coefficients gene dependency set is formed. From every pair of gene dependencies in the gene dependency set, similarity coefficient is measured between two genes using Jaccard Coefficient and thus a gene similarity matrix is computed and a rank is set for each gene indicating its importance. The highest rank gene is considered as the core or the most important gene of the gene set. Next, a rough set theory-based quick reduct algorithm is applied to select only the most informative genes, called reduct, which are sufficient to fully characterise the overall class structure of the gene dataset for disease analysis. Finally, from the reduced gene set of all samples, a rule-based classifier, namely, decision tree is constructed which is applied to unknown samples to predict if it is a diseased or normal sample. Experimental results show the effectiveness of the algorithm.