Proactive fault tolerance for HPC with Xen virtualization
Arun Babu Nagarajan, Frank Mueller, Christian Engelmann, Stephen L. Scott · 2007
Large-scale parallel computing is relying increasingly on clusters with thousands of processors. At such large counts of compute nodes, faults are becoming common place. Current techniques to tolerate faults focus on reactive schemes to recover from faults and generally rely on a checkpoint/restart mechanism. Yet, in today's systems, node failures can often be anticipated by detecting a deteriorating health status.