Incremental checkpoint based failure-aware scheduling algorithm in grid computing
Manjeet Singh · 2016
In Grid computing environment application scheduling is very crucial because the resources are more heterogeneous, geographically distributed, complex and owned by different organizations, they are more prone to failures. Generally, during application/job scheduling only performance factor of resources are considered. But if a node with high computational power also have high failure rate, then there is no such benefit of allocating task to that node because every time a failure occurs it needs recovery and in turn costs in term of time. Thus, failure increases make-span for the job and decreases system/node performance. So, it would be a great idea if we take into consideration failure rate and computational capacity of resources during scheduling. In this paper, to improve the system performance we have proposed a failure-aware scheduling algorithm by taking into consideration both performance and failure factors.