Transparent service reconfiguration after node failures
Thomas Becker · Cooperative Distributed Systems · 1992
The author describes a technique which allows application programs to run in a fault tolerant way. The technique is based on the concept of checkpointing from an active program to one or more passive copies at periodic instances of time. The passive copies serve as an abstraction of stable memory. In case the primary copy fails, exactly one of the backup copies takes over and resumes processing service requests. After each failure a new backup copy is created in order to preserve the fault tolerance degree of the service. A key issue of the technique is the fact that all mechanisms necessary to achieve and maintain fault tolerance can be added automatically to the program code of a non-fault tolerant server, thus making fault tolerance completely transparent for the application programmer.