Statistical and Computational Theory and Methodology for Big Data Analysis
Ming‐Hui Chen, Chuanhai Liu · 2014
The integration of computer technology into science and daily life has enabled the collection of massive volumes of data, such as high-throughput biological assay data, climate data, website transaction logs, and credit card records. However, such big data sets cannot be practically analyzed on a single commodity computer because their sizes are too large to fit in memory or it is too time consuming to process when the current statistical methods are used. To circumvent this obstacle, one may have to resort to parallel and distributed architectures, with multicore and cloud computing platforms providing access to hundreds or thousands of processors. While the parallel and distributed architectures present new capabilities for storage and manipulation of data, from an inferential point of view, it is unclear how the current statistical methodology can be transported to the paradigm of big data. Also, with growing size typically comes a growing complexity of data structures, of the patterns in the data, and of the models needed to account for the patterns. Big data has put a great challenge on the current statistical methodology.