Improving failure diagnosis via better design and analysis of log messages
Ding Yuan · 2012
As software is growing in size and complexity, accompanied by vendors ’ increased time-to-market pressure, it has become increasingly difficult to deliver bulletproof software. Consequently, soft-ware systems still fail in the production environment. Once a failure occurs in production systems, it is important for the vendors to trouble-shoot it as quickly as possible since these failures directly affect the customers. Consequently vendors typically invest significant amounts of resources in production failure diagnosis. Unfortunately, diagnosing these production failures is notoriously difficult. Indeed, constrained by both privacy and expense reasons, software vendors often cannot reproduce such failures. Therefore, support engineers and developers continue to rely on the logs printed by the run time system to diagnose the production failures. Unfortunately, the current failure diagnosis experience with log messages, which is colloquially referred as “printf-debugging”, is far from pleasant. First, such diagnosis requires expert knowl-edge and is also too time-consuming, tedious to narrow down root causes. Second, the ad-hoc nature of the log messages is frequently insufficient for detailed failure diagnosis.