A Backtracking Based Approach for Unsupervised Learning for Mixed Type of Data through K-Mean Based Method
Rohit Rastogi, Arpita Srivastava, Megha Gaur · 2012
Clustering analysis is a descriptive task that seeks to identify homogeneous groups of objects based on the values of their attributes. As we know the advantage of k-means that it converges very fast for forming the clusters and some initial points have to be chosen as seed values and then they are compared with the available elements values and affinity or dissimilarity between them can be calculated by Euclidian/Manhattan/Minkowaski distances. For this purpose K-medoid and K-modes techniques can also be used. The affinity or proximity or dissimilarity is matched with the threshold value for the minimum criteria coverage and if that is met then we say that there is much probability that the new value can lie in that cluster. In this paper, we have tried to design a novel Backtracking based unsupervised learning algorithm which uses K-mean (In future it can be extended to K-mode/K-medoid algorithms) as subroutine to calculate and decide the affinity among data elements and run time dynamism-expansion can be introduced in it by adding the outlier or new uncovered data elements as a new cluster. I. THE CORE IDEA OF THE PAPER As we know that the elements with the least dissimilarity are clustered together. Those elements which do not lie in any cluster, they are treated as outlier in static clustering methods and in dynamic efficient clustering methods, a new cluster can be generated with that unmatched element and that is also treated as initial seed value. This new cluster k is added in the solution set for next new values. Now dynamism in real time application appears here that next incoming values compatibility are checked from first to this k-th new cluster and much probability is that new value is settle down in already available clusters. Very less probability that any exceptional value arrives in the experiment and does not settle down in previously designed clusters so repeat the above said process of defining the new cluster.