FAULT-TOLERANCE OPERATO FOR DISTRIBUTED REAL-TIME CONT

D. Hen · 1990

As distributed computing systems are increasingly used in critical real-time applications such as air-trajflc control, patient monitoring, or power plant control, users' demands for ,fuult-tolerance capabilities of such control computing system are also steadily increasing. In this paper, we are concerned with fault-tolerance at the applicative level i.e. the level of the control programs writien by the application designer. Our goal is to provide high-level architecture independent facilities for allowing control progranis running on a distributed computing system to tolerate failures of the underlying hardware and to tolerate software design faults. Since it is now widely accepted thut systems should be designed and implemented so that the,y are well-structured, we propose a model based on a modular structiirution of the application tasks which permits the building of' dependable control structures b.y use of redundancy. The paper presents an inipletnentation of such an approach to faulttolerance. It is based on a control languuge, Syter, developed in our rescarcli team arid intended to support distributed real-time applications. This language provides various pre-defined operators to express sequential, synchronized, or parallel execution of processes, us well as the preemption and the priority ofsoine evolutions over others. Moreover it offers a set if operators dedicated to frrult-tolerance, The paper describes some of theni .

Read the paper · More papers on PaperTik