Automatic reconfiguration in the presence of failures

Flaviu Cristian · Software Engineering Journal · 1993

We describe a new kind of distributed system service, the availability management service, responsible for ensuring that the critical services of a distributed system remain continuously available to users despite arbitrary numbers of concurrent node removals and node restarts caused by failures, maintenance, and growth. We stress the main ideas behind this new service, and outline a simple design that depends on the existence of synchronous membership and atomic broadcast group communication services. Extensions of this initial design to deal with asynchronous group communication services are also briefly discussed.

Read the paper · More papers on PaperTik