LogPM: Character-Based Log Parser Benchmark

Shayan Hashemi, Jesse Nyyssölä, Mika Mäntylä · 2024

Log parsers transform free-form textual log messages into categorical data and are important tools in automated log analysis pipelines. However, selecting a suitable log parsing algorithm poses a formidable obstacle, thereby underscoring the importance of having a comprehensive benchmark to facilitate decision-making. This paper introduces a novel log parsing benchmark, focusing on predicted template precision at the character level rather than accurate grouping used in the past. We present a new metric called Parameter Mask Agreement that measures template accuracy at the character level alongside a dataset tailored for the task. We identified several challenges that parsers encounter for each dataset, which can aid in developing new log parsers. Moreover, a small empirical study was conducted using the proposed benchmark, evaluating the performance of three renowned parsers: Drain, Spell, and Lenma. The findings revealed that Lenma demonstrated the highest parsing accuracy, whereas Drain exhibited superior parsing speed. Finally, we propose that our benchmark is more appropriate than previous approaches in scenarios where accurate template detection is essential and computational efficiency needs to be assessed.

Read the paper · More papers on PaperTik