On the Survivability of Standard MPI Applications
Anand Tikotekar, Chokchai Leangsuksun, Stephen L. Scott, Box Leangsuksun, Stephan L. Scott · 2006
Job loss due to failure represents a common vulnerability in High Performance Computing (HPC), especially in the Message Passing Interface (MPI) environment. Rollback-recovery has been used to mitigate faulty issues for long running applications. However, to date, the rollback-recovery such as checkpoint mechanism alone may not be sufficient to ensure fault tolerance for MPI applications due to a static view of MPI cooperating machines and lack of resilient ability to endure outages. In fact, MPI applications are prone to cascading failures, where one participating node causes the total failure. In this paper, we address fault issues in the MPI environment by improving runtime availability with self-healing and self-cloning that tolerates the outage of cluster computing systems. We develop a framework that augments a standard HPC cluster with a fault tolerance capability at job level that preserves the job queue, and a parallel MPI job submitted through a resource manager enabling the non-stop execution even after encountering failure.