Feature Selection for Density-Based Clustering

Yun Ling, Chongyi Ye · 2009

In recent years, the advent of high throughput data generation techniques have increased not only the number of objects collected in databases, but also the number of attributes describing these objects. Clustering is the process of grouping the data into classes or clusters, so that objects within a cluster have high similarity in comparison to one another but are very dissimilar to objects in other clusters. Dissimilarities are assessed based on the attribute values describing the objects. Real data are noisy due to measurement technology limitation and experimental variability which prohibits cluster models from revealing true clusters corrupted by noise. In this paper, we utilize correspondence analysis algorithm to process feature selection and then make use of density-based approach for clustering. We find that utilizing the two methods synthetically is very significative to solve actual problems. Experiments on synthetic and real world data demonstrate the efficiency and effectiveness of our algorithm.

Read the paper · More papers on PaperTik