Fault-tolerant parallel matrix multiplication with one iteration fault detection latency
C. Hong, Bruce McMillin · 2002
A new algorithm, the ID algorithm, is presented which minimizes the fault-detection latency. In the ID algorithm, a fault is detected as soon as the fault occurs instead of at problem termination. For n/sup 2/ processors, the fault-latency time of the ID algorithm is 1/n of that of the checksum algorithm with a run-time penalty of O(n log/sub 2/ n) in an n*n matrix operation. This algorithm has better performance in terms of error coverage and expected run time in large-scale matrix multiplications such as signal and image processing, weather prediction, and finite-element analysis.>