Redoop: Supporting Recurring Queries in Hadoop

Chuan Lei, Elke Angelika Rundensteiner, Mohamed Y. Eltabakh · 2014

The growing demand for large-scale data analytics ranging from online advertisement placement, log processing, to fraud detection, has led to the design of highly scalable data-intensive computing infrastructures such as the Hadoop platform. Recurring queries, re-peatedly being executed for long periods of time on rapidly evolv-ing high-volume data, have become a bedrock component in most of these analytic applications. Despite their importance, the plain Hadoop along with its state-of-art extensions lack built-in support for recurring queries. In particular, they lack efficient and scal-able analytics over evolving datasets. In this work, we present the Redoop system, an extension of the Hadoop framework, designed to fill in this void. Redoop supports recurring queries as first-class citizen in Hadoop without sacrificing any of its core features. More importantly, Redoop deploys innovative window-aware opti-mization techniques for recurring query execution including adap-tive window-aware data partitioning, window-aware task schedul-ing, and inter-window caching mechanisms. Redoop retains the fault-tolerance of MapReduce via automatic cache recovery and task re-execution support. Our extensive experimental study with real datasets demonstrates that Redoop achieves significant run-time performance gains of up to 9x speedup compared to the plain Hadoop. 1.

Read the paper · More papers on PaperTik