Automatic service availability management
Flaviu Cristian · 2002
A new kind of distributed system service called availability management service is introduced. It is responsible for ensuring that the critical services of a distributed system remain continuously available to users despite arbitrary numbers of concurrent node removals and node restarts caused by failures, maintenance, and growth. The description of many details involved in a realistic design is sacrificed to make the underlying concepts easily understandable. To this end, the availability management service is designed on top of an easy-to-understand synchronous communication environment, and only one kind of service availability policy is considered. It is indicated how the initial specification and design can be extended to deal with asynchronous systems subject to partitioning as well as with other kinds of service availability policies.>