Clustering for high dimensional data
Varun Kumar Sharma, Anju Bala · 2014
Clustering is an exploratory data analysis technique, which categorizes the dataset into some groups. These groups are formed in a way so that items which have similar features live in same group and those have dissimilar features remain in other. There are many clustering algorithm available. Different kinds of algorithms are best used for different kinds of data. K-means is most used clustering analysis algorithm. It is an iterative approach of point assignment into k clusters. It gives best result and is easily implementable. The k-means algorithm has many issues with it. The main issue is its high time complexity. Several improvements have been suggested by research community. But when it is applied on high dimensional data, the complexity becomes infeasible. In this paper, an approach to reduce the computation of distance function has been proposed. It aims to define a cluster membership set for every cluster. The distance function is calculated only for the clusters which are contained in this set. With this membership set of cluster, the complexity of overall algorithm is reduced.