Coordinated Fault Tolerance for High-Performance Computing
Univ. of Tennessee, Knoxville, TN (United States), Jack J. Dongarra, USDOE Office of Science (SC), George Bosilca · 2013
Our work to meet our goal of end-to-end fault tolerance has focused on two areas: (1) improving fault tolerance in various software currently available and widely used throughout the HEC domain and (2) using fault information exchange and coordination to achieve holistic, systemwide fault tolerance and understanding how to design and implement interfaces for integrating fault tolerance features for multiple layers of the software stack—from the application, math libraries, and programming language runtime to other common system software such as jobs schedulers, resource managers, and monitoring tools.