Automated problem diagnosis in distributed systems
Barton P. Miller, Alexander V. Mirgorodskiy · 2006
Quickly finding the cause of software bugs and performance problems in production environments is a crucial task that currently requires substantial effort of skilled analysts. Our thesis is that many problems can be accurately located with automated techniques that work on unmodified systems and use little application knowledge for diagnosis. This dissertation identifies main obstacles for such diagnosis and presents a three-step approach for addressing them. First, we introduce self-propelled instrumentation, an execution monitoring approach that can be rapidly deployed on demand in a distributed system. We dynamically inject a fragment of code referred to as the agent into one of the application processes. The agent starts propagating through the system, inserting instrumentation ahead of the control flow in the traced process and across process and kernel boundaries. As a result, it obtains the distributed control flow trace of the execution. Second, we introduce a flow-separation algorithm, an approach for identifying concurrent activities in the system. In parallel and distributed environments, applications often perform multiple concurrent activities, possibly on behalf of different users. Traces collected by our agent would interleave events from such activities thus complicating manual examination and automated analysis. Our flow-separation algorithm is able to distinguish events from different activities using little user help. Finally, we introduce an automated root cause analysis approach. We focus on identification of anomalies rather than massive failures, as anomalies are often harder to investigate. Our techniques help the analyst to locate an anomalous flow (e.g., an abnormal request in an e-commerce system or a node of a parallel application) and to identify a function call in that flow that is a likely cause of the anomaly. We evaluated these three steps on a variety of sequential and distributed applications. Our tracing approach proved effective for low-overhead on-demand data collection across the process and kernel boundaries. Manual trace examination enabled us to locate the causes of two performance problems in a multimedia and a GUI application, and a crash in the Linux kernel. Our automated analyses enabled us to find the causes of four problems in the Score and Condor batch scheduling systems.