A distributed fault-tolerant asynchronous algorithm for performing N tasks.

G. Weerasinghe, Lester Lipsky · 2001

This paper is a performance study of a distributed fault-tolerant asynchronous algorithm for performing a job consisting of N independent and idempotent tasks on P processors. The algorithm tolerates up to P 1 processor failures. That is, at least one processor must survive for the lifetime of the application. Processors communicate only through asynchronous message passing. A parameter called Periodicity controls how often progress information is distributed to the rest of the processors. The major design goals are: to eliminate the requirement of a master processor; to optimize the scheduling of tasks such that in the presence of failures and communication time-outs, the number of tasks redone is minimized; to minimize the allocation of resources.

Read the paper · More papers on PaperTik