Parallel Recomputing:A New Approach for Fault-tolerant High Performance Computing
Haifang Zhou · 2009
Checkpointing is the most commonly used scheme for tolerating faults in high-performance computing systems.But this scheme has its performance limitation when the number of processors becomes much larger.The paper proposed a new approach called parallel recomputing for tolerating a single process failure in parallel computing.The main feature of our approach is that it utilizes the computing power of the redundant processor instead of the storage capacity.The paper also presented an optimization of this approach which is a combination of parallel recomputing and checkpointing,and then illustrated how to incorporate parallel recomputing and its optimization into a parallel program.Experimental results demonstrate that the overhead of parallel recomputing is less than checkpointing when the number of processors becomes large,and its optimization can provide a better performance than parallel recomputing.