Analysis of Error Propagation Between Software Processes
Sizarta Sarshar · InTech eBooks · 2011
Nuclear Power -System Simulations and Operation 70•A fault -is a defect within the system.• An error -is a deviation from the required operation of the system or subsystem.•A system failure -occurs when the system fails to perform its required function.This chapter is structured as follows: Section 2 gives a definition of error propagation, describes the mechanisms of error propagation, and previous work on the topic.Section 3 describes the proposed method for studying error propagation between software processes.Section 4 reports on the applicability of the method on one module of the SCORPIO framework.Section 5 addresses the main results.Section 6 discusses the work while section 7 provides conclusions and comments on future work. BackgroundThis section gives a definition of error propagation, describes the mechanisms of error propagation, operating systems and related work on the topic. Error propagationIn our work, error propagation is defined as the situation where an error (or failure) propagates from one entity to another (Sarshar et al., 2007).Errors can propagate between different types of entities, including: physical entities, processes running on single or multiple CPUs, data objects in a database, functions in a program, and statements in a program.Our approach concerns propagation of errors between processes running on a single CPU computer.Systems of interest in our work have not been limited to those that are safety critical only, e.g.systems that are directly involved in controlling a nuclear reactor.A problem of particular interest is the possible negative effect a low criticality application might have on a higher criticality application by means of error propagation because they share common resources.Programs make use of interaction methods provided by the underlying operating system to communicate with each other, or make use of shared resources.These services are provided through the system call interface of the operating system, and are usually wrapped in functions available using standard libraries.Such interaction methods can cause errors and provide mechanisms for error propagation.A coding fault which may be manifested as an error may in principle be anything, e.g. an incorrect instruction or an erroneous data value.It may be manifested inside a local function or an external function.The propagated error need not be of the same type in different functions, e.g. an instruction error in one function realization causes a data error in another.Even if an error is propagated to one function, this does not necessarily mean that the source function fails functionally.The propagated error may only be a side-effect in this function.Another type of error related to function usage is error caused by passing illegal arguments to functions or misusing their return variables.Error propagation between two programs may occur even if both programs individually operate functionally correct.This can e.g.be caused by erroneous side effect in the implementation or execution of the programs.There are two situations possible for how one process can cause another process to fail: • One process experiences a failure, which then causes another process to fail.• One process propagates a fault to another process while not failing itself.According to (Fredriksen & Winther, 2007), possible ways of characterizing error propagation is as either intended or unintended communication or as resource conflicts.www.intechopen.com