An Effective Log Sequence Anomaly Detection Method Guided by Large Language Models

Zhihui Xie, Xiaolin Chai, Hong Yan Luo, Yan Sun · 2025

With the rapid advancement of information technology, the volume of log data generated by information systems has expanded significantly. The effective detection and processing of anomalies in log sequences has emerged as a crucial research topic. Recent studies have successfully applied pre-trained models for anomaly detection in log sequences. However, due to the semi-structured nature of logs and the presence of domain-specific vocabulary, along with the diverse logging styles of different vendors, those models struggled to provide a comprehensive understanding and unified representation of logs. Additionally, current research primarily focused on log template sequences, neglecting the impact of variable values and contextual relationships among logs. In this paper, we propose LlmBertLog, a BERT-based log sequence anomaly detection model guided by large language models. We first utilize a large language model to transform a limited set of raw logs into natural language descriptions as sample data. Subsequently, we induct three novel pre-training tasks designed to enhance BERT's comprehension and representation capabilities through methods such as semantic augmentation and semantic masking. We then integrate a Transformer with positional encoding and fine-tune LlmBertLog for log sequence anomily detection task, thereby improving its performance in log sequence content mining. Extensive experiments demonstrate that LlmBertLog achieves F1 scores of 0.994 on the BGL dataset and 0.998 on the HDFS dataset for log sequence anomaly detection, surpassing the current best benchmarks by 2.3% and 5.5%, respectively.

Read the paper · More papers on PaperTik