Extreme scale and bleeding edge technology lead to a need for resilient high performance computing systems

Nathan DeBardeleben · 2016

High Performance Computing (HPC) and supercomputing are an important sector of the computing field. Distinguishing itself from cloud computing, HPC systems are sized to run extremely large and important calculations such as tightly coupled numerical simulations. These often run for days to weeks on supercomputers and are used to inform scientific discovery and national security. It comes as no surprise then that reliable HPC systems are integral to producing believable scientific results. In this paper we build on years of experience studying supercomputers around the U.S. Department of Energy (DOE) and bring together insights about challenges and needs for HPC reliability. First, we discuss the state of the practice in HPC reliability. We look at a sampling of results from previous work and how reliability telemetry data is used by vendors, system architects, and end users. Finally, we discuss some changes coming in the next decade for HPC systems and how the technology we depend on will drive new reliability challenges that must be addressed and monitored.

Read the paper · More papers on PaperTik