Analysis of Colon Cancer Dataset using K-Means based Algorithms & See5 Algorithms

R. Srinivasa Perumal · 2011

Data mining is used in several medical applications like tumor classification, prediction of medical test effectiveness, genomics, proteomics and DNA sequence analysis. Cancer detection is one of the hot research topics in the bioinformatics age. Data mining techniques, such as pattern association, classification and clustering is applied over gene expression data for detection of cancer. Accuracy is the vital thing to be considered during estimation over colon data. Association works on the basis of correlation, classification helps in categorizing and locate accurately, and clustering is the unsupervised learning ability that is able to discover hidden patterns of dataset. The objective of our work is to make comparative study about various clustering algorithms like simple K-means, global K-means, K-means++ and C5 over cancer dataset is made. Clustering algorithms are compared based on accuracy.

Read the paper · More papers on PaperTik