Cluster analysis on high dimensional data 2010 Presentation
Roland E. Winkler · elib (German Aerospace Center) · 2010
A data set may contain of one or more 'clouds' of data objects. The task for cluster analysis is, to find the location of these clouds and to provide a partitioning of the data objects. Many algorithms are available and often very successful as long as the number of attributes (dimensions) is reasonably low. In high dimensions however, many clustering algorithms do not provide meaningful results any more. In this talk, I will give an overview on the challenges and principle problems of high dimensional data in clustering. The Statistics Department, TU Dortmund provided an artificial example data set that is connected to the BaBar experiment of the Stanford Linear Accelerator. This data set is used as an example to show the problems of high dimensional data sets.