An ai-based architecture of self-stabilizing fault-tolerant distributed process control programs and its analysis

Ing-Ray Chen, Farokh Bastani · 1988

Critical process-control systems must provide uninterrupted services even in the face of partial failures. An architecture based on the notion of self-stabilization and telescopic replication is proposed as a general framework for achieving fault tolerance for these systems. A process-control program is first decomposed into centralized and decentralized subgoals with the objective of simplifying program design. Then, decentralized subgoals are made inherently fault tolerant and assigned to real-time controllers for dealing with stringent real-time control tasks. Centralized subgoals are assigned to controllers which employ AI knowledge representation, planning and learning mechanisms to make the architecture as a whole fault tolerant. We first give a model for describing a process-control system and present a systematic way of decomposing a control program's goal into centralized and decentralized subgoals so as to facilitate organization of control processes in a hierarchical structure. Issues which may arise with decentralized decomposition, along with their possible solutions, are addressed and the notion of inherently fault tolerant programs resulting from decentralized decomposition is defined. For centralized subgoals we investigate AI learning and planning mechanisms to improve the system reliability by means of (a) reducing the strategy formulation time, and (b) reducing the response execution time. In learning, this involves the effect of learning capability on the reliability of the system and how a system can deal with unexpected situations. In planning, this involves the tradeoff analysis between the optimality of a control strategy and the satisfaction of a real-time constraint and the reliability assessment of various search strategies which are used most extensively in AI search paradigm. We also investigate an AI-based knowledge representation method for providing cost-effective fault tolerance for upper level controllers based on the notion of telescopic replication. Our comparative analysis shows that for systems with limited hardware resources (e.g., no repair capability), such as unmanned airspace systems, telescopic replication has a significant advantage over full replication, which is generally used in conventional triplex or duplex systems. A detailed case study of applying telescopic replication to a simulated chemical batch reactor system is shown and the result correlates well with our theoretical prediction.

Read the paper · More papers on PaperTik