APPLICATION-TRANSPARENT ERROR-RECOVERY TECHNIQUES FOR MULTICOMPUTERS †
Tiffany M. Frazier, Yuval Tamir · 1989
We describe and compare error recovery techniques for restoring a valid system state in large ‘‘general-purpose’’ multicomputers following component failures. The techniques discussed are application-transparent, i.e., they do not require application processes to be aware of the fault tolerance characteristics of the system. These techniques are based on checkpointing and rollback: process states are periodically saved to stable storage and are restored from stable storage in the event of a hardware failure. We divide existing recovery schemes into two classes: Message Logging and Coordinated Checkpointing, describe several techniques in each class, and present the advantages and limitations of the schemes when used to provide fault tolerance for large multicomputers. I. Introduction