Two-level checkpoint/restart modeling for GPGPU

Supada Laosooksathit, Nichamon Naksinehaboon, Chokchai Leangsuksan · 2011

Due to the fact that the reliability and availability of a large scaled system inverse to the number of computing elements, fault tolerance has become a major concern in high performance computing (HPC) including a very large system with GPGPU. In this paper, we propose a checkpoint/restart mechanism model which employs two-phase protocol and a latency hiding technique such as CUDA streams in order to achieve a low checkpoint overhead. We introduce GPU checkpoint and restart protocols. Also, we show experimental results and analyze the influences of the mechanism, especially in a long-running application.

Read the paper · More papers on PaperTik