Towards a Statistical Theory of Clustering
Ulrike von Luxburg, Shai Ben-David, Fraunhofer Ipsi · 2005
Abstract. The goal of this paper is to discuss statistical aspects of clus-tering in a framework where the data to be clustered has been sampled from some unknown probability distribution. Firstly, the clustering of the data set should reveal some structure of the underlying data rather than model artifacts due to the random sampling process. Secondly, the more sample points we have, the more reliable the clustering should be. We discuss which methods can and cannot be used to tackle those prob-lems. In particular we argue that generalization bounds as they are used in statistical learning theory of classification are unsuitable in a general clustering framework. We suggest that the main replacements of general-ization bounds should be convergence proofs and stability considerations. This paper should be considered as a road map paper which identifies im-portant questions and potentially fruitful directions for future research about statistical clustering. We do not attempt to present a complete statistical theory of clustering. 1