The Median Resource Failure Checkpointing

Suleman Khan, Khizar Hayat, Sajjad Ahmed Madani, Samee U. Khan, Joanna Kołodziej · 2012

In grid computing, the realization of an enviable fault tolerance ability is linked with the proper utilization of resources and scheduling of jobs. The literature offers two solutions to these two challenging tasks, viz. checkpointing and replication. A checkpointing strategy is being proposed that uses the median of failure intervals of the resources in deciding the checkpoint intervals for the given jobs. The strategy shows improved system throughput, job losses and job execution times while eliminating unnecessary checkpoints.

Read the paper · More papers on PaperTik