Bayesian models for unsupervised feature selection

Yue Guan · 2012

This dissertation focuses on developing probabilistic models for unsupervised feature selection. High-dimensional data often contain irrelevant and redundant features, which can hurt learning algorithms. One can remove these unwanted features either through removing some subsets of the original features (feature selection) or by transforming data into a lower dimensional feature space. Principal component analysis (PCA) is a popular transformation-based dimensionality reduction method. However, it is not easy to interpret which of the original features are important in PCA. We have designed sparse probabilistic PCA and mixture of sparse probabilistic PCA formulations. By presenting sparse PCA as a probabilistic Bayesian formulation, we gain the benefit of automatic model selection. We examined three different priors for achieving sparsification: (1) a two-level hierarchical prior equivalent to a Laplacian distribution and consequently to an L1 regularization, (2) an inverse-Gaussian prior, and (3) a Jeffreys prior. We learn these models by applying variational inference.

Read the paper · More papers on PaperTik