Minimum entropy data partitioning

Stephen John Roberts · 1999

Problems in data analysis often require the unsupervised partitioning of a data set into clusters. Many methods exist for such partitioning but most have the weakness of being model-based (most assuming hyper-ellipsoidal clusters) or computationally infeasible in anything more than a 3dimensional data space. We re-consider the notion of cluster analysis in informationtheoretic terms and show that minimisation of partition entropy can be used to estimate the number and structure of probable data generators. The resultant analyser may be regarded as a Radial-Basis Function classifier. 1 Introduction Many problems in data analysis, especially in signal and image processing, require the unsupervised partitioning of data into a set of `self-similar' clusters or regions. An ideal partition unambiguously assigns each datum to a single cluster and one thinks of the data as being generated by a number of data generators, one for each cluster. Many algorithms have been proposed for such analy...

Read the paper · More papers on PaperTik