A Performance and Energy Comparison of Fault Tolerance Techniques for Exascale Computing Systems
Daniel Dauwe, Sudeep Pasricha, Anthony A. Maciejewski, Howard Jay Siegel · 2016
As the computing power of large scale computing systems increases exponentially with time, their failure rates are increasing exponentially as well. While current high performance computing (HPC) systems experience failures of some type every few days, projections indicate that the next generation exascale machines will experience failures up to several times an hour. The resilience techniques implemented in today's HPC and cloud computing systems do not efficiently scale up to the exascale level. Several more scalable resilience techniques have recently been proposed for next generation HPC systems. However, thus far it is not known how these techniques directly compare to one another in terms of performance and energy use. This work explores four resilience techniques, and considers each technique's ability to handle varying levels of system reliability and system sizes. We demonstrate how each technique compares in terms of application performance and energy use and provide recommendations on their suitability at an exascale level.