Multilevel Clustering for Large Databases

Yves Lechevallier, Antonio Ciampi · Birkhäuser Boston eBooks · 2007

Standard clustering methods do not handle truly large data sets and fail to take into account multilevel data structures. This work outlines an approach to clustering that integrates the Kohonen Self-Organizing Map (SOM) with other clustering methods. Moreover, in order to take into account multilevel structures, a statistical model is proposed, in which a mixture of distributions may have mixing coefficients depending on higher-level variables. Thus, in a first step, the SOM provides a substantial data reduction, whereby a variety of ascending and divisive clustering algorithms becomes accessible. As a second step, statistical modeling provides both a direct means to treat multilevel structures and a framework for model-based clustering. The interplay of these two steps is illustrated on an example of nutritional data from a multicenter study on nutrition and cancer, known as EPIC.

Read the paper · More papers on PaperTik