On communication support for fault tolerant process groups
Ken Birman, Thomas Joseph · 1986
Status of this Memo.This memo describes a collection of multicast communication primitives integrated with a mechanism for handling process failure and recovery.These primitives facilitate the implementation of faulttolerant process groups, which can be used to provide distributed services in an environment subject to non-malicious crash failures.Unlike other process group approaches, such as Cheriton's "host groups" (RFC's 966, 988, [Cheriton]), our approach provides powerful guarantees about the behavior of the communication subsystem when process group membership is changing dynamically, for example due to process or site failures, recoveries, or migration of a process from one site to another.Our approach also addresses delivery ordering issues that arise when multiple clients communicate with a process group concurrently, or a single client transmits multiple multicast messages to a group without pausing to wait until each is received.Moreover, the cost of the approach is low.An implementation is being undertaken at Cornell as part of the ISIS project.Here, we argue that the form of "best effort" reliability provided by host groups may not address the requirements of those researchers who are building fault tolerant software.Our basic premise is that reliable handling of failures, recoveries, and dynamic process migration are important aspects of programming in distributed environments, and that communication support that provides unpredictable behavior in the presence of such events places an unacceptable burden of complexity on higher level application software.This complexity does not arise when using the fault-tolerant process group alternative.This memo summarizes our approach and briefly contrasts it with other process group approaches.For a detailed discussion, together with figures that clarify the details of the approach, readers are referred to the papers cited below.Distribution of this memo is unlimited.