Theoretical Foundations for Clustering and Screening Heterogeneous and High dimensional Data
Yun peng Wei · Deep Blue (University of Michigan) · 2020
This thesis is devoted to the study of two problems in statistics which involve complex data structure of high heterogeneity or large scale. Heterogeneous data arises when each data sample can come from a multiplicity of distributions, creating a population distribution that is a finite mixture of these distributions. The large scale may be due to large data samples, or the large number of variables related to the samples, or the imagined infinite dimensions in nonparametric or functional data. Mixtures of product distributions are a powerful device for learning about heterogeneity within data populations. In this class of latent structure models, de Finetti's mixing measure plays the central role for describing the uncertainty about the latent parameters representing heterogeneity. In the first part of this thesis posterior contraction theorems for de Finetti's mixing measure arising from finite mixtures of product distributions will be established, under the setting the number of exchangeable sequences of observed variables increases while sequence length(s) may be either fixed or varied. The role of both the number of sequences and the sequence lengths will be carefully examined. In order to obtain concrete rates of convergence, a first-order identifiability theory for finite mixture models and a family of sharp inverse bounds for mixtures of product distributions will be developed via a harmonic analysis of such latent structure models. This theory is applicable to broad classes of probability kernels composing the mixture model of product distributions for both continuous and discrete domain $Xfrak$. Examples of interest include the case the probability kernel is only weakly identifiable in the sense of [Ho and Nguyen 2016], the case where the kernel is itself a mixture distribution as in hierarchical models, and the case the kernel may not have a density with respect to a dominating measure on an abstract domain such as Dirichlet processes. An important problem in large scale inference is the identification of variables that have large correlations or partial correlations with at least one other variable. Recent work in correlation screening has yielded breakthroughs in the ultra-high dimensional setting when the sample size $n$ is fixed and the dimension $p rightarrow infty$ (see [Hero and Rajaratnam 2012]). Despite these advances, the correlation screening framework suffers from some serious practical, methodological and theoretical deficiencies. For instance, theoretical safeguards for partial correlation screening requires that the population covariance matrix be block diagonal. This block sparsity assumption is however highly restrictive in numerous practical applications. As a second example, results for correlation and partial correlation screening framework requires the estimation of dependence measures or functionals, which can be highly prohibitive computationally, rendering the framework impractical and unappealing in the very setting it is designed for. In the second part of this thesis, we propose a unifying approach to correlation and partial correlation screening which specifically goes beyond the block diagonal correlation structure, thus yielding a methodology that is suitable for modern applications. By making insightful connections to random geometric graphs, the number of highly correlated or partial correlated variables are shown to have a novel compound Poisson limit, and are obtained for both the finite $p$ case and when $p rightarrow infty$. The unifying framework also demonstrates an important duality between correlation and partial correlation screening with important theoretical and practical consequences.