A Survey Of Recoverable Distributed Shared Memory Systems
Christine Morin, Isabelle Puaut, Centre National de la Recherche Scientifique (CNRS), 35 - Rennes (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA), Rennes-1 Univ., 35 (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA), Institut National des Sciences Appliquees de Rennes (INSA), 35 (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA), Institut National de Recherche en Informatique et en Automatique (INRIA), 35 - Rennes (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA) · 1995
Distributed Shared Memory (dsm) systems provide a shared memory abstraction on distributed memory architectures (distributed memory multicomputers, networks of workstations). Such systems ease parallel application programming since the shared memory programming model is often more natural than the message-passing paradigm. However, the probability of failure of a dsm system increases with the number of sites. Thus, fault tolerance mechanisms must be implemented in order to allow processes to continue their execution in the event of a failure. This paper gives an overview of recoverable dsm systems (rdsm) that provide a checkpointing mechanism to restart parallel computations, after a site failure.