A State-Machine Approach to Disambiguating Supercomputer Event Logs

Jon Stearley, Robert A. Ballance, Lara E Bauman · 2012

Supercomputer components are inherently stateful and interdependent, so accurate assessment of an event on one component often requires knowledge of previous events on that component or others. Administrators who daily monitor and interact with the system generally possess sufficient operational context to accurately interpret events, but researchers with only historical logs are at risk for incorrect conclusions. To address this risk, we present a state-machine approach for tracing context in event logs, a flexible implementation in Splunk, and an example of its use to disambiguate a frequently occurring event type on an extreme-scale supercomputer. Specifically, of 70,126 heartbeat stop events over three months, we identify only 2% as indicating failures. as reboots during scheduled maintenance) from the data. Our results show that such consideration can dramatically affect event counts, and thus subsequent reliability metrics. Several researchers have applied state-machine models to the analysis of event logs. Song used three states (up, down, warm) in a Markov model of system availability, and applied it to logs from the White supercomputer (7). A state machine was also applied to Hadoop logs for the purpose of debugging programs based on anomalous time periods between state transitions (8). These studies use small models for very specific purposes, and do not address operations context such as

Read the paper · More papers on PaperTik