Fault-Tolerant Execution of Computationally and Storage Intensive Parallel Programs Over a Network of Workstations: A Case Study

Judith Smith, Santosh Kumar Shrivastava · 1995

This paper addresses the question as to whether there is potential gain to be made from executing such computations in such an environment. The example application considered here is matrix multiplication. As well as being data intensive, this is one of the class of embarrassingly parallel applications which do not require synchronization between concurrent processes during the course of the computation. A simple analysis, backed up by experimental results, shows what sort of speedup may be expected from the experimental configuration. The network used consists of HP710 and HP730 workstations connected to a 10 Mbps ethernet and is used in this establishment for general purpose computing. As the scale of a distributed computation is increased in either the number of participating nodes or its duration, the possibility of a failure occurring which might affect the execution of the computation must increase. For example, the owner of a workstation which is participating in a computation may choose to reboot his machine. If it is not possible to tolerate such an event, it is necessary to restart the entire computation. Whether the computation continues after the fault or is restarted, application data structures on disk should be made consistent. The second aspect of the work described here attempts to address such issues of fault tolerance through a fault-tolerant implementation of the well known bag of tasks structure, which is described in [7]. The fault tolerance is achieved through the use of atomic actions (equivalent here to atomic transactions)

Read the paper · More papers on PaperTik