Global events and global breakpoints in distributed systems
Dieter Haban, W. Weigel · 2003
A solution to the problem of setting breakpoints in distributed systems is described. It is shown what kind of breakpoints are possible, how to detect those breakpoints, and how to halt the system in a consistent state. The communication among the processes may be asynchronous with an arbitrary ordering of messages. The algorithms select between simultaneous events and those ordered according to L. Lamport's happened-before relation (1978). The mechanisms for definition and detection of global breakpoints are implemented in the distributed debugging system DTM (distributed test methodology).>