14. Model-Based Clustering Algorithms

Society for Industrial and Applied Mathematics eBooks · 2007

Model-based clustering is a major approach to clustering analysis. This chapter introduces model-based clustering algorithms. First, we present an overview of model-based clustering. then, we introduce Gaussian mixture models, model-based agglomerative hierarchical clustering, and the expectation-maximization (EM) algorithm. Finally, we introduce model-based clustering and two model-based clustering algorithms.14.1 IntroductionClustering algorithms can also be developed based on probability models, such as the finite mixture model for probability densities. The word model is usually used to represent the type of constraints and geometric properties of the covariance matrices (Martinez and Martinez, 2005). In the family of model-based clustering algorithms, one uses certain models for clusters and tries to optimize the fit between the data and the models. In the model-based clustering approach, the data are viewed as coming from a mixture of probability distributions, each of which represents a different cluster. In other words, in model-based clustering, it is assumed that the data are generated by a mixture of probability distributions in which each component represents a different cluster. Thus a particular clustering method can be expected to work well when the data conform to the model.Model-based clustering has a long history. A survey of cluster analysis in a probabilistic and inferential framework was presented by Bock (1996). Early work on model-based clustering can be found in (Edwards and Cavalli-Sforza, 1965), (Day, 1969), (Wolfe, 1970), (Scott and Symons, 1971b), and (Binder, 1978). Some issues in cluster analysis, such as the number of clusters, are discussed in (McLachlan and Basford, 1988), (Banfield and Raftery, 1993), (McLachlan and Peel, 2000), (Everitt et al., 2001), and (Fraley and Raftery, 2002).

Read the paper · More papers on PaperTik