Improving availability in distributed systems with failure informers
Joshua B. Leners, Trinabh Gupta, Marcos K. Aguilera, Michael Walfish · 2013
Abstract. This paper addresses a core question in dis-tributed systems: how should applications be notified of failures? When a distributed system acts on failure re-ports, the system’s correctness and availability depend on the granularity and semantics of those reports. The system’s availability also depends on coverage (failures are reported), accuracy (reports are justified), and time-liness (reports come quickly). This paper describes Pi-geon, a failure reporting service designed to enable high availability in the applications that use it. Pigeon exposes a new abstraction, called a failure informer, which al-lows applications to take informed, application-specific recovery actions, and which encapsulates uncertainty, al-lowing applications to proceed safely in the presence of doubt. Pigeon also significantly improves over the pre-vious state of the art in the three-way trade-off among coverage, accuracy, and timeliness. 1