Failure Detection in Asynchronous Distributed Systems
Raimundo José de Araújo, Macêdo, Campus de Ondina · 2000
Being able to detect failures is an important issue in designing fault-tolerant distributed systems. However, the actual behaviour of a system limits the ability to provide such a mechanism. From one e xtreme of the spectrum, synchronous s ystems (i.e., with bounded message transmission delay and processing times) allow for the construction of perfect failure detection based simply on local timeouts. At t he other extreme, accurate failure detection cannot be developed for asynchronous s ystems (i.e. systems with no bounds on message transmission delays and processing times), unless some e xtra properties can be guaranteed, such the ones specified in a seminal article by Chandra a nd Toueg [1]. The present paper discusses the requirements and describes the implementations of f ailure detectors for two important fault-tolerant mechanisms meant t o asynchronous environments: process group membership and S Failure Detector based d istributed consensus [1]. These implementations are based on a mechanism called the Time Connectivity Indicator, introduced in this paper.