Fault tolerant distributed computing system for heterogeneous network: Application to non-destructive testing
Vivek Ramdas Gadiyar · 1996
Rapid technological advances in the field of communications and microelectronics has led to increased interest in distributed computing systems. Researchers have focused their attention towards issues concerning structure of hardware and design of software to exploit the inherent parallelism and fault tolerance in distributed systems. Due to the existence of diverse hardware and software platforms, heterogeneity in distributed systems cannot be avoided. Heterogeneous environments present new challenges because they introduce incompatibility due to differing communication protocols, operating systems, file systems and databases. This work addresses the problems mentioned above and presents a distributed computing system that is fault tolerant and works on a heterogeneous network of workstations. The systems utilizes the enormous processing power that is available on a network of workstation to execute compute-intensive tasks. It is possible to achieve tremendous speed-up if the compute-intensive task can be broken into units that can be executed independently. Among the various computationally intensive applications this thesis identifies Non-destructive evaluation techniques as a medium of illustrating the computational power of distributed systems. Further, non-destructive evaluation techniques involve simulations that exhibit coarse grain parallelism, which can be efficiently exploited in distributed systems. Parallel environments are often organized as multiple processing units with a single functional block to synchronize these processing units. The processing units are delegated the subtasks by a scheduling system. A loss of a subtask in such a parallel computing environment results in permanent blocking of the module which performs synchronization. This work provides mechanisms to identify and restart lost subtasks to achieve fault tolerant operation. The subtasks are distributed over idle nodes on the network to achieve maximum speed-up and efficiency. This strategy also increases the throughput of the network due to maximum utilization of the idle nodes on the network.