Robust clustering algorithm for high dimensional data classification based on multiple supports
Benson S. Y. Lam, Hong Yan · 2008
High dimensionality, noisy features and outliers can cause problems in cluster analysis. Many existing methods can handle one of the problems well but not the others. In this paper, we propose a new clustering algorithm to solve these problems. The basic idea is to control the support of the optimization procedure so that the effect produced by those contaminated samples and dimensions is greatly reduced. This is achieved by using multiple supports. Initially, a large support is used and then its size is reduced and eventually only a subgroup of data samples is considered for clustering. This procedure can filter out lots of contaminated information. Experiment results show that the proposed method effectively resolves all these problems. It outperforms existing ones for real world high dimensional datasets.