Data Labeling method for genome DNA data based on Cluster similarity using Rough Entropy for Categorical Data Clustering

Mr.Sreenivasulu G, Dr.Viswanadha Raju S, Sambasiva Rao N · International Journal of Engineering and Technology · 2017

Clustering is one of the major issues in data mining.Data labeling has been recognized as an important method in categorical clustering.Clustering is technique where all similar data point are grouped.However, with data labeling is applied on those points which are not labeled earlier.Although there are many approaches in the numerical domain, but very limited algorithms are available for categorical data.To address this problem of how to allocate those unlabeled data points into proper clusters remains as a challenging issue in the categorical domain.In this paper, a mechanism is proposed for labeling and keeping the similar data points into accurate clusters.We have a data set named Genome DNA where grouping of 'superfluous' Splice junctions on those points on a DNA sequence is a major challenge.The predicament posed in this dataset is to recognize, given a sequence of DNA, the limits between exons and introns.The new proposal is to allocate each unlabeled data point into the equivalent proper cluster with data labeling also.This method has two advantages: 1) The proposed method exhibits high execution efficiency.2) This method can achieve quality clusters.The proposed method is empirically validated on DNA data set, and it is shown significantly more efficient than prior works while attaining results of high quality.

Read the paper · More papers on PaperTik