Data Integration and Model Fusion in the Bayesian and Frequentist Frameworks

Emily C. Hector, Lu Tang, Ling Zhou, Peter X.‐K. Song · 2024

Current statistical practice is greatly challenged by rapid advances in technologies that enable practitioners to frequently collect massive and complex data to study various problems of practical importance. Scalability, studied and implemented in the divide-and-conquer framework, is a key feature of statistical methods developed to address the computational needs of modern data analysis [ 27 ]. Divide-and-conquer strategies generally proceed in three steps: first, the whole data is divided into smaller and more manageable subsets, or blocks, of data; then each data block is analyzed in parallel; finally, results from data block analyses are combined under some optimality criterion. In contrast, many practical studies, such as genetic studies, attempt to combine datasets collected under similar study protocols to form a mega dataset in order to increase the sample size and variety of features for high statistical power and generalizability. In both cases, the essential statistical difficulty concerns data integration and model fusion.

Read the paper · More papers on PaperTik