A framework for analyzing software system log files
Mladen Alan Vouk, Meiyappan Nagappan · 2011
It is the premise of this work that current log analysis methods are too ad hoc and do not scale well enough to be effective in the domain of large logs (such as those we might expect in a computational cloud system). The more complex the system, the more complex and voluminous its logs are. In this dissertation we investigate, identify and develop components needed for an adaptable end-to-end framework for the analysis of logs. The framework needs to take into consideration that different users look for different kinds of information in the same log files. Required are adaptable techniques and algorithms for efficient and accurate log data collection, log abstraction and log transformations. The techniques or algorithms that are used in each component of the framework will vary according to the application, its logging mechanisms and the information that the stake holder needs to make decisions. We discuss a concrete instantiation of the framework (i.e. specific solutions for each component) and explain the context in which it is used. Throughout the dissertation we have used the log files from the cloud computing infrastructure at North Carolina State University, called the 'Virtual Computing Laboratory' (VCL) (see Appendix A for more details). We built the operational profile of VCL from its log files, and the developers used this for regression testing, feature selection, and performance improvement. We were successful in identifying the most frequent and least frequent set of events (operational profile) by using two different analysis techniques. Our approach involves transforming the log files using either suffix array data structure or a weighted directed cyclic graph data structure. The current state of the art in building operational profiles from log files offers only semi automated or manual approaches. Using these two transformations, we were able to get the actual usage frequency of VCL use cases, in a automated, linearly scaling approach. In order to build these transformations we had to abstract the semi structured log messages to specific log events. The current state of the art in abstracting semi structured logs involves either using regular expressions, which takes a long time, but provides accurate results, or using heuristic approaches that are much faster but provided poorer accuracy. We propose an empirical approach, in addition to the heuristic approach, to maintain the speed but improve the accuracy. We were able to reduce the inaccuracies in abstraction by half (down from 5% to 3%), as compared to the pure heuristic based approaches, when we used a hybrid approach. In the case of large logs, this can represent a considerable improvement. For example, in the VCL case this improved the abstraction accuracy of about 16,000 events in a daily log file. On one hand, we need more data to be more accurate, on the other hand we need less data to make analysis faster. Unfortunately, common current log collection mechanisms can either collect a detailed level of logs at all times or basic level of logs at all times. Ideally, one would collect more data when necessary and less when not required. We propose such a technique. It uses a trained decision engine to intelligently collect information from the application to the log file. By using this technique we were able to reduce the size of the log by almost 29% and still retain the data that the developers want to inspect. This represents a major contribution to the state-of-the-art in accurate but efficient log processing. (Abstract shortened by UMI.)