Classification of Cancer Types based on Gene Expression Data
Yinchao He, Ryan Bockmon, Miracle Modey, Sarah Roscoe · 2020
With the rapid increase of genomic data, analyzing this huge bioinformatics data becomes a new challenge. Therefore, computer-based processes and algorithms are much more powerful to use. For analysis of across tumor types, John N Weinstein declares the uncertain feasibility of applying characterization based on molecular changes to complement pathological analysis for classification of cancers. Hence, for this limitation, we compared four deep learning methods, K-Means, Support Vector Machine, Formal Concept Analysis, and Association rules, to classify cancers based on the gene expression cancer RNA-Seq data set with 801 cases and 20531 genes for each case. The results show that SVM has the highest accuracy (99.8% and 99.2%) out of all our objectives, which is followed by K-Means with 91.75%. The third highest overall result (83.1%) is the Formal Concept Analysis algorithm. Association rules have the lowest accuracy with 72.25%. This comparison supplies a good guide for the classification of cancer types based RNA-seq.