Software Implemented Fault Tolerance: Technologies and Experience
Yennun Huang, Chandra M. R. Kintala · 1993
By software implemented fault tolerance, we mean a set of software facilities to detect ‘and recover from faults that are are not handled by the underlying hardware or operating system. We consider those faults that cause an application process to crash or hang; they include software faults as well as faults in the underlying hardware and operating system layers if they are undetected in those layers. We define 4 levels of software fault tolerance based on availability and data consistency of an application in the presence of such faults. Watchd, libft and nDFS are reusable components that provide up to the 3rd level of software fault tolerance. They perform, respectively, automatic detection and restart of failed processes, periodic checkpointing and recovery of critical volatile data, and replication and synchronization of persistent data in an application software system. These modules have been ported to a number of UNIX’ platforms and can be used by any application with minimal programming egort. Some newer telecommunications products in AT&T have already enhanced their fault-tolerance capability using these three components. Experience with those products to date indicates that these modules provide eficient and economical means to increase the level of fault tolerance in a software product. The performance overhead due to these components depends on the level and varies from 0.1% to 14% based on the amount of critical data being checkpointed and replicated.