Proactive blocking coordinated checkpointing with dynamic intervals

Mehdi Lotfi, Seyed Ahmad Motamedi, Mojtaba Bandarabadi · 2009

In this paper we introduce a new proactive blocking coordinated checkpointing for cluster computing systems with dynamic interval. Many current schemes to increase the availability of cluster computing systems either make use of redundancy in space or redundancy in time (reactive methods). These methods induce the overhead to the cluster computing system in failure free execution time. In order to minimize the performance loss (rollback and checkpoint overheads) due to unexpected failures or unnecessary overhead of fault tolerant mechanisms, we present a proactive method for the blocking coordinated checkpointing strategy. Existing checkpointing methods are static with constant checkpointing interval. These methods are based on the exponential distribution function. In this paper we use the Weibull distribution function to find the dynamic interval. Our method is based on the failure data analysis of LANL cluster system. Experimental results show that average execution time of NAS application is significantly reduced by using the proposed method.

Read the paper · More papers on PaperTik