Validation of a dynamic checkpoint mechanism for Apache Hadoop with failure scenarios
Paulo Vinícius Cardoso, Patrícia Pitthan Barcelos · 2018
New computational paradigms have created data intensive applications which have a demand for efficient and reliable processing platforms. High performance systems, used to answer this demand, have a increasing number of components such as nodes and cores. A multi component system may suffer with reliability and availability issues once the mean time between failures become smaller. Checkpoint and Recovery (CR) is a fault tolerance technique based on backward error recovery that focus on retrieving system safety state from backup saves. This paper shows the Checkpoint and Recovery technique implemented by Apache Hadoop, a framework that allows distributed processing of large datasets across clusters of computers. Hadoop uses the checkpoint technique to provides fault tolerance on Hadoop Distributed File System (HDFS). However, choosing an appropriate checkpoint interval is a major challenge once Hadoop defines the CR attributes statically. Then we propose a dynamic solution for checkpoint attributes configuration on HDFS, whose goal is to make it adaptable to system usage context. We expose a validation of both static and dynamic mechanisms on failure induced scenarios with DataNode crashes in order to determine the overhead of checkpoint and recovery steps.