Failure-aware scheduling in grid considering Weibull failure distribution

Manjeet Singh, Ritu Garg · 2013

In Grid computing environment, as the resources are more heterogeneous, geographically distributed, complex and owned by different organizations, they are more prone to failures. Application scheduling in such a environment is very crucial. Generally, during application/job scheduling only performance factor of resources are considered. But if a node with high computational power also have high failure rate, then there is no such benefit of allocating task to that node because every time a failure occurs it needs recovery and in turn costs in term of time. Thus, failure increases make-span for the job and decreases system/node performance. A node with comparatively lower computational capacity and lower failure rate may give better performance and reduced make-span. So it would be a great idea if we take into consideration failure rate and computational capacity of resources during scheduling. In this paper we have proposed an approach for scheduling the tasks. We recalculate the computational capacity of resources by finding the expected wasted time due to presence of failure and then the tasks are scheduled according to this new computational capacity. Here the failure of nodes is treated as to follow Weibull distribution.

Read the paper · More papers on PaperTik