Failure and its Recovery in an Object-Oriented Distributed System

Stephen Crane, Brendan Tangney · 1991

This paper describes a method for recovering permanent object state in an object-oriented distributed system. Inspiration for this work was derived from observation of the lengths to which programmers have traditionally been forced to go in order to make their programs resilient to failure. This experience led to the decision that such a burden was unacceptable and that the onus of recovery be shifted onto the underlying operating system. Further goals were that the user be insulated to the greatest possible degree from failure and its recovery (failure transparency) and that the resulting system be as efficient as possible under normal conditions. 1 Author's current address is: Imperial College, Department of Computing, 180 Queen's Gate, London SW7 2BZ. 1 Introduction In a distributed system linked by a network of some kind, failure is partial. That is, a single failure affects only a subset of the available equipment. This is in contrast to centralised (or stand-alone) systems in...

Read the paper · More papers on PaperTik