Proactive Fault Tolerance in Large Systems

Sayantan Chakravorty, Celso Luiz Mendes, Laxmikant V. Kalé · 2004

High-performance systems with thousands of processors have been introduced in the recent past, and systems with hundreds of thousands of processors should become avail-able in the near future. Since failures are likely to be fre-quent in such systems, schemes for dealing with faults are important. In this paper, we introduce a new fault tolerance solution for parallel applications that proactively migrates execution from a processor where a failure is imminent. Our approach assumes that some failures are predictable, and leverages the fact that current hardware devices contain various fea-tures supporting early indication of faults. By using the con-cepts of processor virtualization in Charm++ and Adaptive MPI (AMPI), we describe a mechanism that migrates ob-jects when a failure is expected to arise in a given proces-sor, without requiring spare processors. After migrating ob-jects, and applying a load balancing scheme, the execution of an MPI application can proceed and achieve optimized efficiency. We modify the implementation of collective oper-ations, such as reductions, so that they continue to operate efficiently even after a processor is evacuated and crashes. To demonstrate the feasibility of our approach, we present preliminary performance data. 1

Read the paper · More papers on PaperTik