TRCID: Optimized Task Recovery in MapReduce Based on Checkpointing Intermediate Data
Peng Wang, Jing Liu, Kai Ding · 2017
The computing paradigm of MapReduce is popular for data analysis since the volume of data to analyze grows rapidly. Hadoop, an open-source implementation of MapReduce, is the most popular framework for MapReduce programs. But it still suffers from the job delay caused by node failure and task failure. This paper proposes a novel Task Recovery mechanism based on Checkpointing Intermediate Data (TRCID) to improve the performance of Hadoop when failure happens. We keep track of the intermediate data to make it available to be used directly or regenerated efficiently after the recovery. Rather than waiting for reducers fetch the output of mappers, we send intermediate data proactively on regular intervals. At the same time the meta data will be replicated to the distributed file system as a checkpoint file. Upon failure detection, the task can be resumed according to the checkpoint instead of being restarted from the beginning. Compare with original Hadoop, TRCID enhances the fault tolerant capability for MapReduce framework, and has better computing performance.