A FRAMEWORK FOR ADAPTIVE FAULT MANAGEMENT IN DISTRIBUTED COMPUTING SYSTEMS
Rex E. Gantenbein, Thomas F. Lawrence, Sung Y. Shin · International Journal of Reliability Quality and Safety Engineering · 1994
Most strategies for fault management in distributed computing systems are effective only for a narrow range of fault classes. In some systems, however, a wide range of operating environments may be encountered that require different strategies to be used at different times. Adaptive fault management can be used to increase a system’s survivability under conditions that can suddenly and drastically change, but there are a number of problems that must be solved for this paradigm to be successful. Two of these problems are the characterization of fault management techniques and the evaluation of system behavior relative to system requirements. A taxonomy for distributed fault management characterizes a variety of well-known techniques for achieving survivability in distributed systems. A generalized metric can then evaluate the effectiveness of different fault management techniques under changing conditions. These two concepts may be used as a framework for understanding how an adaptive system can dynamically select an appropriate fault management strategy from a number of alternatives in response to changes in its operating environment. Applications that use multiple fault management techniques to enhance their survivability illustrate the applicability of this framework.