Observation and Control for Debugging Distributed Computations

Vijay K. Garg · 1997

I present a general framework for observing and controlling a distributed computation and its applications to distributed debugging. Algorithms for observation are useful in distributed debugging to stop a distributed program under certain undesirable global conditions. I present the main ideas required for developing efficient algorithms for observation. Algorithms for control are useful in debugging to restrict the behavior of the distributed program to suspicious executions. It is also useful when a programmer wants to test a distributed program under certain conditions. I present different models and their limitations for controlling distributed computations. 1 Introduction Many problems in distributed systems, especially in distributed debugging, can be viewed as special cases of the problem of observing and controlling a distributed computation. For example, deadlock detection, termination detection, and breakpoint detection are some specific instances of the observation problem...

Read the paper · More papers on PaperTik