Generic Faults and Architecture Design Considerations in Flight-Critical Systems
Stephen S. Osder · Journal of Guidance Control and Dynamics · 1983
Specific examples of generic faults involving interaction of hardware and software in flight-critical systems are described. Such faults can cause an avalanche shutdown of all redundant computer channels. The popular technique of redundant computer frame synchronization is shown to be particularly vulnerable. Architecture solutions that allow dissimilar redundancy and incorporate brick-wall isolation are described. Practical techniques of coping with time-skew effects in unsynchronized computer channels are given, and it is shown that they offer many simplicity advantages. HE term ''generic became part of the technical language to describe design defects that elude the test and analysis procedures used to validate a redundant control system design. Although the existence of such defects can be postulated in any type of system, the generic fault concept is especially significant in the flight-critical system application because it defeats the massive redundancy strategies that designers rely on to meet safety or reliability objectives. A vision of such a fault toppling each channel of a flight-critical redundant system is a nightmare that haunts the designer and the certifier communities. The first part of this paper examines various classes of such faults, acknowledging that there can be no 100% certainty that a system is free of them. Specific illustrations show why their detection and correction are difficult. It is contended that redundancy architecture is a key factor in generic fault vulnerability, with redundant channel syn- chronization as a major aggravator of error mechanisms. Well-known solutions to such vulnerabilities involve concepts of channel decoupling, including the brick wall approaches that were popular in early analog mechanizations.1 '2 The term brick wall refers to extraordinary measures taken to insure the physical and electrical separation of redundant channels. Dissimilar redundancy also has its advocates,3 but both solutions preclude the advantages provided by frame (or loose) synchronization of computer channels. Such syn- chronization allows an inherent simplification of monitoring algorithms. Various degrees of synchronizati on tightness have been used. The Space Shuttle computers synchronize to sublevels of frames or tasks, 4 while some systems under development use complete microclock synchronization in massively redundant central processing unit (CPU) and memory structures.5'6 When channel decoupling and isolation (including dissimilar hardware and software mechanisms) are used to avoid generic-fault vulnerabilities, channel syn- chronization must then be abandoned. A recurring theme of this paper is that the synchronization of redundant computers opens pathways for generic faults to cause multiple channel shutdowns. In some instances these faults reside in incomplete logic designs that might not an- ticipate the multiplicity of computer resynchronization in- teractions following transient events in the power distribution system. In other instances, the generic fault could also exist in the nonsynchronized systems, but an event which triggers the fault would cause only one channel to shut down, whereas the same fault in synchronous architectures would cause multiple