Design and implementation of an availability management service

S. Mishra, Guozhao Pang · 2003

This paper describes the design and implementation of an automatic availability management service called Teams for a timed asynchronous distributed system. Teams automatically reconfigures a distributed system in the presence of communication and node failures in such a way that all computing services remain available, and the system reconfiguration is transparent to the users. Teams is a fail-aware service: a node at any point in time knows whether it can provide the automatic reconfiguration services or not. It can reconfigure a system in response to node failures, communication partitions, or maintenance operations in less than 8 milliseconds.

Read the paper · More papers on PaperTik