AN ENHANCED APPROACH FOR PROJECTING CLUSTERS IN HIGH DIMENSIONAL SPACES
Arun Shalin · 2012
Clustering high-dimensional data has been a major challenge due to the inherent sparsity of the points. Most existing clustering algorithms become substantially inefficient if the required similarity measure is computed between data points in the full- dimensional space. To address this problem, a number of projected clustering algorithms have been proposed. However, most of them encounter difficulties when clusters hide in subspaces with very low dimensionality. To overcome the difficulties, a robust partitioned distance-based projected clustering algorithm (PCKA) is presented in the previous work (1) describes the process of identifying the low dimensional values in high dimensional space and avoids the computation of the distance in the full dimensional space. But the existing PCKA algorithm analyzed only about the attribute relevance and does not discussed the analysis of redundancy in the database. To enhance the process of projecting clusters of high dimensional subspaces, an enhanced framework is presented in this work which solves the projected clustering problem. Even though the number of dimensions for all the clusters specific subspace varies, the process of identifying the single small subset of dimensions for all the clusters is achieved efficiently. An experimental evaluation is carried out to estimate the performance of the proposed enhanced approach for projecting clusters in high dimensional spaces in terms of average cluster dimensionality, outlier immunity and consumption time.