P−data integration
Kim‐Anh Lê Cao, Zoe Marie Welham · 2021
Biological findings from high-throughput experiments often suffer from poor reproducibility. This is most likely due to the small number of samples inherent in high-dimensional data. One way to improve statistical power and the reproducibility of results is to combine data sets from independent experiments measured on the same P variables in an integrative analysis. However, study, or batch effects must be considered when integrating independent studies to enable genuine biological variation to be identified. This chapter describes a recently developed method in mixOmics based on multi-group PLS-DA to integrate independent studies while taking study effects into account. The chapter describes the tuning criteria for key input arguments, the various types of numerical and graphical outputs, and criteria for model assessment. The case study stemcells available from mixOmics includes gene expression data sets from four independent studies. The chapter concludes by introducing other examples of P− integration for microbiome and single cell transcriptomics data.