Abstractions for Fault-Tolerance.
Flaviu Cristian · 1994
ions for Fault-Tolerance Flaviu Cristian, Computer Science and Engineering University of California, San Diego, CA 92093-0114 Designing and understanding fault-tolerant distributed system architectures is notoriously difficult: one has to maintain control not only over standard (failure-free) behaviors, but also over a multitude of failure behaviors caused by component failures. The lack of clear structuring concepts and terminology can exacerbate this difficulty. This paper complements earlier attempts at introducing some order and discipline in this area [5], by discussing a number of basic concepts and services that simplify the understanding and design of fault-tolerant systems. Fault-tolerance has two different meanings. First, a system is said fault-tolerant if its behavior remains well-defined when components fail. For example, a storage service that either reads correctly a value written previously or signals an exception is fault-tolerant in the above sense: low level bit corr...