Computer event monitoring and analysis

Michael F. Buckley · 1992

Availability, which is defined as the proportion of time that a system is available to perform its specified function, is an important attribute of computer systems. It can be improved by using Symptom Directed Diagnosis (SDD). This thesis uses the largest set of event logs studied to date to examine three issues related to SDD. These are event monitoring, the distribution of the time between events and event clustering. This research includes the most exhaustive investigation of the monitoring process conducted to date. The results show that higher quality event logs are needed. Deficiencies in the monitoring process complicate the analyses, consume time and make incorrect assumptions and wrong conclusions more likely. The impact of the deficiencies on the results depends upon the particular analysis. One example shows that inappropriate handling of events with incorrect dates produces an order to magnitude change in the mean Time To Tuple (TTT). Recommendations for improving the monitoring process are given. The Time Between Reboot (TBRb) and Time Between Crash (TBCr) variables were compared to five distribution models, namely the normal, lognormal, exponential, Weibull and gamma. The results show that none of the five models are a good match for either variable. There is evidence to suggest that the lack of fit is due to the variables being summations of a number of distributions rather than single distributions. Tupling rules are heuristics used to group events. The rules proposed by Tsao (1983) and expanded by Hansen (1988) were applied to the event logs. The results were combined with knowledge of the monitoring process to obtain a semantic understanding of why the heuristics work. A number of improvements for the VAX/VMS logs are proposed to increase the effectiveness of tupling. Descriptive univariate statistics for the rules are also presented. The results indicate that tupling is a useful and general methodology for performing lower level associations between events. The high degree of skew in the tuple entry count and tuple span variables indicates that a simple thresholding scheme based on the 95th or 99th percentile would be an effective alarm mechanism.

Read the paper · More papers on PaperTik