Distributed fault-tolerance for large multiprocessor systems

Jon G. Kuhl, Sudhakar M. Reddy · 1980

Techniques for dealing with hardware failures in very large networks of distributed processing elements are presented. A concept known as distributed fault-tolerance is introduced. A model of a large multiprocessor system is developed and techniques, based on this model, are given by which each processing element can correctly diagnose failures in all other processing elements in the system. The effect of varying system interconnection structures upon the extent and efficiency of the diagnosis process is discussed, and illustrated with an example of an actual system.

Read the paper · More papers on PaperTik