Feature Selection with Efficient Initialization of Clusters Centers for High Dimensional Data Clustering
Dharmveer Singh Rajput, Pramod Kumar Singh, Mahua Bhattacharya · 2011
Most of the traditional data clustering algorithms suffer from two main problems (i) the curse of dimensionality and (ii) random initialization of clusters centers which leads to local optimum clustering. In this paper, we propose a technique for selecting most relevant dimensions of data set and efficient initialization of clusters centers. Our proposed technique uses the median absolute deviation (MAD) for selection of relevant dimensions of data set and then it uses most frequent value (MODE) of selected dimensions to determine the initial clusters centers. Finally these initial clusters centers are used in the k-means algorithm for optimum clustering. Empirical results show that the algorithm produces comparatively efficient results. The quality measures also validate good quality of the obtained results.