Latent Variable Models and Data Visualisation
Chris Bishop, Michael E. Tipping · 2000
Abstract Visualisation is a powerful and widely used technique for data analysis and data mining. For simple datasets a single projection of the data on to a two-dimensional plane, such as that provided by principal component analysis, may prove adequate. In the case of more complex datasets, however, it may be necessary to find multiple plots corresponding to different projection directions and/or different subsets of the data points in order to capture the full complexity of the data. Here we use latent variable models to construct a framework for data visualisation which allows simultaneous soft clustering and projection of the data in a probabilistic setting. We first show how standard principal component analysis can be formulated in terms of maximum likelihood under a latent variable model. Next we extend the formalism to include both mixtures and hierarchical mixtures of principal component models, and derive the corresponding visualisation algorithms. Finally, we illustrate the hierarchical approach to visualisation using datasets obtained from multiphase flows along oil pipelines, and from satellite image data.