Automated Log Message Embeddings
Adrian Murphy, Daniel Larsson, Fanny Söderlund, Ola Angelsmark, Johan Eker · 2024
System logs are crucial for understanding the state and health of systems, yet manual inspection becomes impractical due to the high volume of messages. Consequently, machine learning-based log anomaly detection has emerged to automatically identify irregularities. This study investigates the effectiveness of log message embeddings, a novel parsing method, for anomaly detection in complex systems. Specifically, we evaluate their resilience to concept drift compared to traditional parsing approaches. The study conducts empirical analyses on benchmark datasets, revealing that log message embeddings achieve comparable anomaly detection results while demonstrating greater robustness against concept drift than traditional methods like Drain. Additionally, the study highlights the usefulness of large language models in automating the log embedding pipeline to handle out-of-vocabulary words and extract synonymous and antonymous relationships. Insights gained from the study suggest potential refinements for future research in this area, contributing to advancements in system monitoring and log anomaly detection.