Performance analysis of checkpoint based efficient failure-aware scheduling algorithm

Manjeet Singh · 2017

As the resources in Grid computing environment are more heterogeneous, geographically distributed, complex and owned by different organizations they are more prone to failures and hence in Grid computing environment, application scheduling is very crucial. There are various techniques used to deal and recover from the failure. One of the most efficient and usable approach is Check pointing. Generally, during application/job scheduling only performance factor of resources are considered (example speed). But if a node with high speed also have high failure rate, then there is no much benefit of allocating task to that node because every time a failure occurs it needs recovery and in turn costs in term of time. Thus, failure increases make-span for the job and decreases system/node performance. So, using the failure information of the node with its performance factors, as a key criterion for scheduling may change the scenario and results in a efficient fault tolerant scheduling algorithm. In this paper, I have analyzed the performance of a failure-aware scheduling algorithm [31] over several performance factors : Performance Ratio, Job Failure Rate, Throughput, Service Unit Loss and Failure Slow Down.

Read the paper · More papers on PaperTik