More than Manuscripts: Reproducibility, Rigor, and Research Productivity in the Big Data Era
Lance Allyn Waller, Gary W. Miller · Toxicological Sciences · 2016
Data Scientist (n.): Person who is better at statistics than any software engineer and better at software engineering than any statistician. —Tweeted by @Josh_Wills The rapid advances in novel measurement technologies, fast processing algorithms, distributed computing, and cloud applications all lead to scientific inquiry based on observations defining the “3 Vs” of big data: volume (big), velocity (fast computing allows fast measurement), and variety (data linked from multiple sources). An increasing number of high-profile studies not only involve data from the authors’ experiments but also relate to data obtained from multiple sources, processed together in an analytic pipeline. The time and effort involved in identifying data sources, obtaining relevant data elements, linking data from disparate sources, and curating the final analytic data set (“data wrangling” or “data munging”) is considerable, and can often be difficult to recreate. In light of recent scandals and close examinations of the reproducibility of high-profile science in an age of big data, the U.S. National Institutes of Health recently announced requirements for “rigor and reproducibility” in funded research (https://www.nih.gov/research-training/rigor-reproducibility).