Probabilistic Models in Partitional Cluster Analysis
Hans Henrich Bock · 2003
Cluster analysis is designed for partitioning a set of objects into homogeneous classes by using observed data which carry information on the mutual similarity or dissimilarity of objects. Clustering methods are often defined in a heuristic or algorithmic way, emphasizing computational aspects and heuristic motivations . In contrast, this paper considers the clustering problem in a probabilistic framework and presents a survey on probabilistic models for partition-type clustering structures . It is shown how clustering criteria and grouping methods may be derived from these models in the case of vectorvalued data, dissimilarity matrices and similarity relations . 1 The clustering problem and its underlying data The ability to classify objects into homogeneous classes on the basis of their mutual similarities, dissimilarities or analogies is a basic element of human intelligence, an undispensible tool for the recognition of visual and conceptual structures, and an indispensable element for any abstract way of thinking . In the framework of statistics and data analysis, the classification problem occurs typically when large sets of objects are described by huge amounts of data which can never be analyzed without a preliminary step of information compression just by detecting or constructing a sufficiently small number of homogeneous classes of (similarly behaving) objects whose properties can be summarized by suitable class prototypes or class-specific feature combinations which provide an easy insight into, and a concise overview of, the full set of data. Or when it is conjectured that a given set of observations originates from several sources and seems to show some obvious heterogeneities : then the revelation of separate classes will reveal the hidden (possibly : causal) data structure and allow the development of class-specific strategies for solving substance-related questions, such as class-specific therapies for patients, group-specific publicity campaigns for attracting consumers, characterizing distinct types of social of psychological behaviour, distinguishing different types of soils or agricultural regions, locating single outlier cases etc. The classification of employees into different salary groups or the segmentation of the clients of an insurance company into types with their distinct risk structure provide examples with a more organizational motivation . 'Institute of Statistics, Technical University of Aachen, D-52056 Aachen, Germany