Scalable group-based checkpoint/restart for large-scale message-passing systems
Justin C. Y. Ho, Cho-Li Wang, Francis C. M. Lau · Proceedings - IEEE International Parallel and Distributed Processing Symposium · 2008
The ever increasing number of processors used in parallel computers is making fault tolerance support in large-scale parallel systems more and more important. We discuss the inadequacies of existing system-level checkpointing solutions for message-passing applications as the system scales up. We analyze the coordination cost and blocking behavior of two current MPI implementations with checkpointing support. A group-based solution combining coordinated checkpointing and message logging is then proposed. Experiment results demonstrate its better performance and scalability than LAM/MPI and MPICH-VCL. To assist group formation, a method to analyze the communication behaviors of the application is proposed.