Reliable job execution with process failure recovery in computational grid
P. Latchoumy, P. Sheik Abdul Khader · International Journal of Information and Communication Technology · 2015
Grid computing provides a virtual framework that integrates heterogeneous resources and services distributed across multiple controlled domains. Dynamic grid environment makes the grid infrastructure unreliable, resulting in failure of executing jobs. The great challenge here is to provide the reliable job execution in the presence of resource failure. In this paper, we present a new model called reliable job execution with process failure recovery in computational grid. In this model, the reliability of sites is monitored and the historical data is used for predicting resource failures taken into account when dispatched user jobs to resources and it recovers the failed job after the process failure has occurred. If the process failure occurs due to CPU overloads or memory thrashing, the backup process starts to run from the recently saved checkpoint. This system also considers the availability of checkpoints by storing checkpoints in multiple backup checkpoint servers. The experimental results demonstrate that our proposed strategy provides the guaranteed service to the grid user within the specified deadline.