Relevance-Based Data Selection for BERT-Based Anomaly Detection Using Unstructured Logs

2025

As software systems grow in scale and complexity, log-based automated anomaly detection has become an essential tool for maintaining system reliability.However, machine learning and deep learningbased detection methods typically require pre-labeled data for training, which poses a challenge due to the vast volume and repetitive nature of logs generated by large systems.Furthermore, log formats often evolve with system updates, making traditional log parsers prone to errors that can negatively affect anomaly detection performance.To address these challenges, this study proposes a robust anomaly detection system that directly processes unstructured logs.The system employs a BERT tokenizer for tokenization and utilizes relevancebased selection and clustering techniques to extract less than 0.01% of high-quality training data from millions of unlabeled logs.Additionally, BERT is leveraged to capture both sequential and semantic information in the logs, facilitating the automated detection of normal and anomalous patterns.Experimental results demonstrate that the proposed method achieves an F1-score exceeding 0.96 across four supercomputer datasets, and an F1score above 0.91 for the detection of previously unseen events.

Read the paper · More papers on PaperTik