Towards High Availability for High-Performance Computing System Services: Accomplishments and Limitations

Christian Engelmann, Steven L. Scott, Leangsuksun Chokchai, Xubin He · 2006

During the last several years, our teams at Oak Ridge National Laboratory, Louisiana Tech University, and Tennessee Technological University focused on efficient redundancy strategies for head and service nodes of high-performance computing (HPC) systems in order to pave the way for high availability (HA) in HPC. These nodes typically run critical HPC system services, like job and resource management, and represent sin-gle points of failure and control for an entire HPC sys-tem. The overarching goal of our research is to pro-vide high-level reliability, availability, and serviceabil-ity (RAS) for HPC systems by combining HA and HPC technology. This paper summarizes our accomplish-ments, such as developed concepts and implemented proof-of-concept prototypes, and describes existing lim-itations, such as performance issues, which need to be dealt with for production-type deployment.

Read the paper · More papers on PaperTik