Performing Replay in an OSF DCE Environment
Yuh Ming Yong, DAVID J. TAYLOR · 1995
Debugging a distributed application is inherently difficult because such factors as network delayandvarying system loads may cause the behaviour of the application to change from one execution to another. Using a standard debugger with such an application is also likely to perturb its behaviour sufficiently that some bugs will not be manifested when the debugger is used. Fortunately, a replay mechanism can be used to circumvent these problems. Replaying a program involves a monitoring phase, during which the behaviour of the program is logged, and a replay phase, during whichthe behaviour of the program is coerced to follow the partial order logged during the monitoring phase. In this paper, we describe an implementation of the replay technique for use with the Open Software Foundation's Distributed Computing Environment (OSF DCE). OSF DCE presents a special problem in that servers are multithreaded and incoming RPCs are assigned dynamically to threads. To preservetheevent partial order, it is necessary to ensure that an RPC uses the same thread on replay. This difficulty is dealt with by ensuring that the DCE thread service sees the same resource state during replay and hence makes the same thread-allocation decision.