Clustering for binary data and mixture models—choice of the model

Mohamed Nadif, G. Govaert · Applied Stochastic Models and Data Analysis · 1997

When cluster analysis is based on mixture models, choosing an appropriate model is a difficult problem. Previous studies usually addressed a part of this problem by estimating the number of clusters and assuming the type of model to be known. Various criteria to be minimized have been proposed to measure a model's suitability by balancing model fit and model complexity. In this work, we extend the work of Govaert (1990) and Celeux and Govaert (1995) to the use of some of these information criteria in the detection of the type of Bernoulli mixture model while assuming that the number of clusters is known. We simulated samples with various underlying types of model and separations of components using Monte Carlo simulations. These simulations show the advantages and the weaknesses of the considered information criteria with a view to determining the type of model. In addition, they underline the importance of a judicious choice of model type in order to obtain a good clustering. © 1998 John Wiley & Sons, Ltd.

Read the paper · More papers on PaperTik