Coordinated Checkpointing using Vector Timestamp in Grid Computing.
Taichi Jinno, Tokimasa Kamiya, M. Nagata · 2006
Abstract In grid computing, system recovery is carried out using checkpoints recorded at each nodes. The resource manager must recover system with keeping global consistency to prevent Domino effect. Currently, coordinated checkpointing is widely used in which all processes can be synchronized. Considering overhead due to synchronization, we will present a coordinated checkpoint protocol using vector timestamp to reduce overhead. Our proposed protocol aims to reduce idle time of every process by grasping occurred event numbers. We will also evaluate performance of the proposed protocol. Experiment was carried out for parallel computation of eight nodes. As the result of the experiment, we obtained reduction of overhead time with 55 percentages in average at each processes. Thus, we showed effectiveness of our proposed protocol for scalable grid computing.