Anomaly Detection in Large-Scale Monitoring Systems using a Language Model
Supasate Vorathammathorn, Nopphakorn Subsa-Ard, Tawan Thaepprasit, Phond Phunchongharn, Sansiri Tarnpradab · 2025
Anomaly Detection in large-scale monitoring systems, especially within high-performance computing (HPC), is a significant challenge because disruptions when the computer running can operations and reduce overall efficiency.We propose a novel framework called Anomaly Detection in Large-Scale Monitoring Systems using a Language Model (AD-LM).This framework uses a language-model-driven workflow for anomaly detection.Starting with, AD-LM applies BERTopic for topic modelling, which groups log entries into meaningful clusters, helping to expose patterns that indicate potential anomalies.The next technique, using a graphbased classification model identifies system failures by capturing key relationships within both HPC and large-scale logs.This framework supports high-speed processing and minimal memory usage-essential qualities in HPC settings.We evaluated AD-LM on three real-world log datasets (Hadoop Distributed File Systems, BlueGene/L, and Thunderbird) and achieved F1-scores of 0.995, 0.997, and 0.998, respectively-outstanding well-known anomalydetection benchmarks with little overhead.Our findings confirm AD-LM's effectiveness for real-time anomaly detection in HPC and large-scale scenarios, underscoring its robustness, adaptability, and efficient resource consumption.