CloneHadoop: Process Cloning to Reduce Hadoop's Long Tail
Sarthak Kukreti, Frank Mueller · 2018
Recent advances in distributed computing have enabled large-scale data processing on high volumes of data with MapReduce. However, the overall runtime of such applications is dependent on the slowest subtask at the tail end of the work queue. Current approaches mitigate this effect by scheduling redundant speculative attempts of straggler tasks. However, speculative attempts have to recompute work already done by the original task. This often prevents late speculations from completing before the straggling task. This work promotes a novel speculation approach via process cloning to avoid redundant computations transparent to users. But speculation requires consistency between task attempts and results in divergent execution of subtasks. The work contributes (1) a process cloning approach for speculative execution, (2) mechanisms to maintain consistency and recover complex components, and (3) an integration of cloning and recovery into Apache Hadoop with optimizations to alleviate resource bottlenecks. In experiments, cloning benefits straggling tasks with runtime reductions of up to 25% on jobs with larger tasks. Our solution scales with different task and problem sizes and is robust across different runtime distributions of straggler tasks.