High performance techniques for reliable execution of tasks under hardware and software faults

Osama Ahmed Abulnaja · 1996

Various studies have shown that both hardware and software are subject to failures. Furthermore, in the highly reliable computing environment, techniques have been developed mostly for dealing with faults and reliable execution of tasks. Generally, in this case system resources have been used inefficiently and no special attempt has been made to optimize the system performance concurrently. On the other hand, in the high performance computing environment, it is often assumed that hardware and software are fault-free and efficient techniques have been developed for fast execution of tasks. In this work, we have studied the development of high performance techniques for reliable execution of tasks and concurrent diagnosis of faults. Our approach is intended to maximize both performance and reliability. We have also evaluated system performance and reliability under the devised techniques. Specifically, we studied the following problems. First, we have studied software fault-tolerant techniques. Motivated by minimizing the effect of redundancy on the system performance, we have introduced a model of computation for improving software reliability and system performance and also have develop new software fault-tolerant scheduling algorithms. In addition, we have evaluated the system performance under the devised software fault-tolerant techniques. System reliability and performance are evaluated by analytical and simulation techniques. Next, we have examined hardware fault-tolerant techniques and have proposed a new model of computation for dealing with hardware faults. The proposed model is devised for reliable execution of tasks and on-line diagnosis of processors and communication channels failures while attempting to concurrently maximize the system performance. New scheduling algorithms based on the proposed model have been introduced and simulation results presented for some of the proposed scheduling algorithms. We have also evaluated a lower bound for a task reliability under the proposed model of computation. Finally, we have considered an integrated approach for dealing with both software and hardware fault-tolerance concurrently. Here, an approach has been proposed for the ultrareliable execution of tasks where both hardware and software are subject to failures and, importantly, on-line fault diagnosis is achieved while concurrently maximizing system performance. Based on this proposed integrated approach, new scheduling algorithms have been introduced and a lower bound for a task reliability has been evaluated.

Read the paper · More papers on PaperTik