Hephaestus: Data Reuse for Accelerating Scientific Discovery
Jennie Duggan, Michael L. Brodie · 2015
Data-intensive science, wherein domain experts use big data analytics in the course of their research, is becoming increas-ingly common in the physical and social sciences. Moreover, data reuse is becoming the new normal, owing to the open data movement [15] and arrival of big science experiments such as the Large Hadron Collider. Here, a small group of researchers with exotic equipment produce a dataset that is shared by thousands. Unfortunately, weak and spurious cor-relations are also on the rise in research [5, 27]. For example, Google Flu Trends published their algorithms in 2008 [19] for use in public health, and in the intervening time its accu-racy has plummeted. In the 2011-2012 flu season, this system produced estimates more than 50 % higher than the number of cases reported by the U.S. Center for Disease Control [32]. This work first examines common pitfalls associated with data-intensive science and how they contribute to irrepro-ducible results. We then propose a system for conducting virtual experiments over existing data. It simulates random-ized controlled trials by reframing the principles of empir-ical research. These virtual experiments underpin a larger platform we call Hephaestus. This framework accumulates virtual experiments in a visualization to help scientists iden-tify consistencies and anomalies in an area of research. We then highlight a set of research challenges associated with this platform. We argue that by using this approach, data-intensive science may come to achieve accuracy on par with its causality-driven predecessors. 1.