Fail‐safe concurrency in theEcliPSesystem

Felipe Knop, Vernon J. Rego, Vaidy S Sunderam · Concurrency Practice and Experience · 1996

Local or wide-area heterogeneous workstation clusters are relatively cheap and highly effective, though inherently unstable operating environments for long-running distributed computations. We found this to be the case in early experiments with a prototype of the EcliPSe system, a software toolkit for replicative applications on heterogeneous workstation clusters. Hardware or network failures in computations that executed for over a day were not uncommon. In this work, a variety of features for the incorporation of failure resilience in the EcliPSe system are described. Key characteristics of this fault-tolerant system are ease of use, low state-saving cost, system scalability and good performance. We present results of some experiments demonstrating low state-saving overheads and small system-recovery times, as a function of the amount of state saved.

Read the paper · More papers on PaperTik