Transparent In-memory Cache for Hadoop-MapReduce
Venkatesh Nandakumar · TSpace (University of Toronto) · 2014
Many analytic applications built on Hadoop ecosystem have a propensity to iteratively perform repetitive operations on same input data. To remove the burden of these repetitive operations, new frameworks for MapReduce have been introduced, which make users follow its programming model. We propose a solution to the problem of application rewriting that newer frameworks impose. We re-architected Hadoop core to add in-memory caching and cache-aware task-scheduling.We set out to match the performance of a state-of-the-art high speed, in-memory MapReduce architecture with caching (Spark). While Spark reimplements the MapReduce paradigm, it comes with a new set of new API's and abstractions. We maintain the familiar Hadoop framework and API's, thus complete backward compatibility for any existing Hadoop-based software. This ensures no changes to existing applications code whatsoever. It guarantees no-pain installation over existing deployments while providing 4.5-12X performance improvement. We perform comparable to, and in some cases outperform Spark.