Dependability evaluation of parallel/distributed computer networks

C.R. Das · 1986

As a result of the proliferation of low cost off-the-shelf microprocessors and the advances in networking technology, parallel/distributed computing has become increasingly popular. Parallel systems are broadly divided into two categories depending on the type of interconnection topology. These are called systems and systems. At the system level, a multiprocessor or multicomputer consists of two subsystems. One subsystem is the computation facility which is provided by processors and memories. The second subsystem is the communication network used to support interprocessor communication. These two components are equally important in a parallel processing environment. Any failure of a component in these subsystems results in performance degradation of the system. This dissertation deals with the dependability evaluation of multiprocessor and multicomputer architectures and their behavior with graceful degradation. Three dependability measures, known as reliability, performance availability (PA), and maintainability, are used to evaluate and characterize various parallel architectures. These measures are based on the system requirements for the parallel execution of a task (job) which consists of a number of subtasks. This research effort differs from the most previous ones in that methods for incorporating the degradation of the communication network into system model are addressed. In the sequel, we present two different models; one is known as the bus oriented model (BOM) and the second one is known as the switch oriented model (SOM), for analyzing the reliability and bandwidth availability (BA) of various multiprocessors. The BOM is an analytical model whereas the SOM is a simulation technique. The SOM captures the communication network representation more accurately. A simulation model is presented to evaluate the reliability and computation-communication availability (CCA) of all types of multicomputer networks suggested in the literature. Mathematical models are presented for the on-line and off-line maintenance of both types of parallel systems. First, the effect of on-line maintenance processor reliability on the system dependability is addressed. Second, models for system downtime and cost are developed for three types of off-line maintenance policies known as scheduled maintenance (SM), unscheduled maintenance (UM), and scheduled and unscheduled maintenance (SUM).

Read the paper · More papers on PaperTik