Proactive fault tolerance for HPC with Xen virtualization

Arun Babu Nagarajan, Frank Mueller, Christian Engelmann, Stephen L. Scott · 2007

Large-scale parallel computing is relying increasingly on clusters with thousands of processors. At such large counts of compute nodes, faults are becoming common place. Current techniques to tolerate faults focus on reactive schemes to recover from faults and generally rely on a checkpoint/restart mechanism. Yet, in today's systems, node failures can often be anticipated by detecting a deteriorating health status.

Read the paper · More papers on PaperTik