A Graph-based Data Mining Approach for Template Recognition using Large Log Datasets in Software Systems

Claudia Crespo Julio, Ahmed Shaharyar Khwaja, Isaac Woungang, Alagan Anpalagan · 2024

Log files generated by software systems can be utilized as a valuable resource in data-driven approaches to improve the system health and stability. These files often contain valuable information about runtime execution, and their effective monitoring requires analyzing an increasingly large volume of data logs. In this paper, a graph mining technique for log parsing is presented, which is source agnostic to the system. This means that the technique can function regardless of the source of the logs, making it more scalable and reusable. Unlike the existing approaches that rely heavily on domain knowledge and regular expression patterns, the proposed approach uses graph models and semantic analysis to detect data patterns with minimal user input. This makes it easy to implement it in a variety of scenarios where application-based logs may differ significantly. The proposed parsing technique is evaluated over seven datasets. It achieves the best performance on the Thunderbird dataset, where the technique takes 3.87 seconds for 2000 logs, while obtaining precision, recall and F1 measure higher than 0.99.

Read the paper · More papers on PaperTik