Fault detection using abstracted models of finite-state machines

Konstantinos Nikolaou Oikonomou · 1983

One approach to fault tolerance in computer systems is the use of one machine to monitor the behavior of another. From a general viewpoint, the problem is to detect faults in a complex system when the observations possible on it and the observer's knowledge of its correct behavior are limited. This problem is studied here by representing the monitored system S by a finite-state machine (fsm) M = (X, Y, Q, (delta), (lamda)). Faults in S are modelled by changes in M's next-state and output maps (delta) and (lamda), which result into M' = (X, Y, Q, (delta)', (lamda)'). The observer (machine) is connected to S through known--but arbitrary--input, output, and state encoders, which let it distinguish only blocks of partitions (PROD)(,X), (PROD)(,Y), (PROD)(,Q) on the sets X, Y, and Q of M. This triple of partitions, called an A, limits the observations possible on S. The observer's knowledge of the correct behavior M of S is correspondingly limited, consisting in a non-deterministic fsm M(,A) = ((PROD)(,X), (PROD)(,Y), (PROD)(,Q), (delta)(,A), (lamda)(,A)). The inputs to M may be (thought of as) generated by a random process known to the observer. In that case, the observer's knowledge of M is represented by a stochastic sequential machine M(,A) obtained by attaching probabilities to the non-deterministic maps (delta)(,A) and (lamda)(,A) of M(,A). The following questions arise in this setting, and are examined in^this thesis. (Q1) Investigate the fundamental limits to fault detection^in an fsm M observed* through A, and whose correct behavior is^known as M(,A) (or M(,A)). Faults in M may or may not affect M(,A) (or M(,A)),^and there are some that are undetectable by any fault detection^scheme. (Q2) Given a complex system model M and a desired^abstraction ratio (a function of (VBAR)X(VBAR) / (VBAR)(PROD)(,X)(VBAR), (VBAR)Y(VBAR) / (VBAR)(PROD)(,Y)(VBAR), (VBAR)Q(VBAR) / (VBAR)(PROD)(,Q)(VBAR)),^find an optimal abstraction A* for fault detection. Subsidiary^questions are: (a) What fault detection scheme should be used? A^statistical procedure is needed in the case of M(,A). (b) What is an^optimal abstraction A*? We introduce the sizes of various^detectable and undetectable fault classes as optimality criteria. (Q3) Extend modelling and fault detection results to systems of interconnected fsm's, representing distributed computer systems. ^^*Note: the observer is not allowed to apply test inputs to M.

Read the paper · More papers on PaperTik