Reducing Faulty Executions of Distributed Systems
Robert Colin Butler Scott · eScholarship (California Digital Library) · 2016
When confronted with a buggy execution of a distributed system—which are commonplacefor distributed systems software—understanding what went wrong requiressignificant expertise, time, and luck. As the first step towards fixing the underlying bug,software developers typically start debugging by manually separating out events that areresponsible for triggering the bug (signal) from those that are extraneous (noise).In this thesis, we investigate whether it is possible to automate this separation process.Our aim is to reduce time and effort spent on troubleshooting, and we do so byeliminating events from buggy executions that are not causally related to the bug, ideallyproducing a “minimal causal sequence” (MCS) of triggering events.We show that the general problem of execution minimization is intractable, but wedevelop, formalize, and empirically validate a set of heuristics—for both partially instrumentedcode, and completely instrumented code—that prove effective at reducing executionsize to within a factor of 4.6X of minimal within a bounded time budget of 12 hours.To validate our heuristics, we relay our experiences applying our execution reductiontools to 7 different open source distributed systems.