A User-triggered Checkpointing Library for Computationintensive Applications.

Geert Deconinck, Johan Vounckx, Rudy Lauwereins, Jean Peperstraete · 1995

: We propose a method to incorporate coordinated checkpointing and rollback in high performance computing applications on massively parallel computers. A library allows the user to specify which data-items (including files) belong to the contents of the checkpoint, and to trigger the checkpointing in the application. The recovery-line management on the distributed disk system takes care of which recovery-lines are valid, (and may be used for rollback), and of which are obsolete (and may be deleted). This flexible approach provides a frame to incorporate non-blocking, continuebefore -validate checkpointing in the application. Besides, the main advantages are the hardware independence and flexibility. Keywords: fault tolerance, checkpointing, backward error recovery, massively parallel systems 1. Introduction Many scientific and commercial applications (simulations, solutions of computational complex problems, ...) require long execution times: they run for days or weeks before the re...

Read the paper · More papers on PaperTik