Cooperative decision-based fault diagnosis
Sampath Rangarajan, Donald S. Fussell · 1990
The need for very reliable computing resources has led to the design of fault tolerant parallel and distributed systems. One of the main requirements for fault tolerance is the ability to diagnose and repair failures in the components of these systems. Much work has been done in providing self diagnosis capabilities for these systems, where components of the system cooperate to make diagnosis decisions on each other. These decisions are unreliable because the components that make these decisions may have failed. In this dissertation, we study diagnosis methods based on cooperative decision making. We propose a general model for these problems where a decision maker determines the status of the components in the system. The components are modeled as events which can take one of two values--faulty(0) or fault-free(1). The interaction between the decision maker and the events is modeled as a communication process where the set of events (sender) communicates with the decision maker (receiver) through a noisy communication channel. The noise in the channel reflects the unreliability of the decisions made. We derive lower bounds on the information transfer from the sender to the receiver in order to determine the values of the events with high probability. Further, we give structure to the decision maker and assume that it is made up of entities each of which is capable of extracting exactly one bit of information from the events. With this assumption, we design simple threshold based voting algorithms based on what we call a HVT model. These algorithms are shown to be optimal in terms of the number of entities that are required. We use the above algorithms for solving three fault diagnosis problems. First, we consider the problem of diagnosing failures in the processors of a multiprocessor system. Our optimal solutions for this problem are the first which do not pose any restrictions on the connectivity of the network connecting the processors in the system. Secondly, we consider the problem of rectifying corrupted files in a distributed file system. Our solutions for this problem are again optimal and improve on existing results. Thirdly, we consider the problem of diagnosing failures in a wafer-scale environment and argue that our solutions for multiprocessor fault diagnosis could be used in such an environment.