Bayesian mixtures for large scale inference

Francesco Denti · RERO DOC (Universität Bern) · 2020

Bayesian mixture models are ubiquitous in statistics due to their simplicity and flexibility and can be easily employed in a wide variety of contexts. In this dissertation, we aim at providing a few contributions to current Bayesian data analysis methods, often motivated by research questions from biological applications. In particular, we focus on the development of novel Bayesian mixture models, typically in a nonparametric setting, to improve and extend active research areas that involve large-scale data: the modeling of nested data, multiple hypothesis testing, and dimensionality reduction. Therefore, our goal is twofold: to develop robust statistical methods motivated by a solid theoretical background, and to propose efficient, scalable and tractable algorithms for their applications. The thesis is organized as follows. In Chapter 1 we shortly review the methodological background and discuss the necessary concepts that belong to the different areas that we will contribute to with this dissertation. In Chapter 2 we propose a Common Atoms model (CAM) for nested datasets, which overcomes the limitations of the nested Dirichlet Process, as discussed in Camerlenghi et al.,2018. We derive its theoretical properties and develop a slice sampler for nested data to obtain an efficient algorithm for posterior simulation. We then embed the model in a Rounded Mixture of Gaussian kernels framework to apply our method to an abundance table from a microbiome study. In Chapter 3 we develop a BNP version of the two-group model (Efron, 2004), modeling both the null density f_0 and the alternative density f_1 with Pitman-Yor process mixture models. We propose to fix the two discount parameters sigma_0 and sigma_1 so that sigma_0>sigma_1, according to the rationale that the null PY should be closer to its base measure (appropriately chosen to be a standard Gaussian base measure), while the alternative PY should have fewer constraints. To induce separation, we employ a non-local prior (Johnson and Rossell, 2010) on the location parameter of the base measure of the PY placed on f_1. We show how the model performs in different scenarios and apply this methodology to a microbiome dataset. Chapter 4 presents a second proposal for the two-group model. Here, we make use of non-local distributions to model the alternative density directly in the likelihood formulation. We propose both a parametric and a nonparametric formulation of the model. We provide a theoretical justification for the adoption of this approach and, after comparing the performance of our model with several competitors, we present three applications on real, publicly available genomic datasets. In Chapter 5 we focus on improving the model for intrinsic dimensions (IDs) estimation discussed in Allegra et al.,2019. In particular, the authors estimate the IDs modeling the ratio of the distances from a point to its first and second nearest neighbors (NNs). First, we propose to include more suitable priors in their parametric, finite mixture model. Then, we extend the existing theoretical methodology by deriving closed-form distributions for the ratios of distances from a point to two NNs of generic order. We propose a simple Dirichlet process mixture model, where we exploit the novel theoretical results to extract more information from the data. The chapter is then concluded with simulation studies and the application to real data. Finally, Chapter 6 presents the future directions and concludes.

Read the paper · More papers on PaperTik