Classification non-supervisée de données de grande dimension et de graphes à l'aide de modèles à variables latentes discrètes
Nicolas Jouvin · HAL (Le Centre pour la Communication Scientifique Directe) · 2020
This thesis proposes three original contributions for the clustering of particular types of data: multivariate continuous and count data, and also networks. First, a new algorithm for the clustering of high-dimensional count data is described, showing a real advantage over competing approaches, especially in small sample size settings. A medical application is detailed, with the clustering of anatomopathological text reports from Institut Curie hospital. Then, a new Gaussian mixture model for high-dimensional continuous data is presented, along with a clustering algorithm relying on unsupervised linear discriminant analysis. The latter compares favorably to state-of-the-art approaches on simulated and real-data benchmarks, and potential extensions are discussed. We finish by proposing a two-fold methodology for hierarchical clustering, based on an exact version of the integrated classification likelihood (ICL). The first part consists in improving existing greedy heuristics, using a carefully designed genetic algorithm to reduce sensitivity to local maxima of the exact ICL. Then, we consider a new asymptotic approximation of the latter giving rise to a hierarchical strategy, merging the clusters obtained from the first algorithm. We show how this approach is generically applicable in any discrete latent variable model for which exact ICL are tractable, and we detail derivations for standard ones. Simulations and real-datasets applications demonstrate the interest of this methodology, in particular for statistical network analysis.