Power Microservices Troubleshooting by Pretrained Language Model with Multi-source Data
Shuai Dominique Ding, Yifei Xu, Zhuang Lu, Fan Tang, Tong Li, Jingguo Ge · 2024
Microservice has become the mainstream paradigm for developing cloud-native applications, but the intricate interdependencies between microservices and the vast amount of heterogeneous observable data (i.e. metrics, logs and traces) pose challenges for rapid troubleshooting. Several anomaly detection and root cause localization approaches that integrate multi-source data have been proposed. However, they are plagued with issues such as scarcity of high-quality data and insufficient model generalization. This is particularly evident when domain-specific models are trained from scratch for specific tasks. Recently, Large Language Models (LLMs) have shown outstanding capabilities in time series analysis, due to multi-source data generated by distributed microservices exhibit intrinsic spatio-temporal characteristics. In view of this, we propose LLM4MST, an LLM-empowered microservice troubleshooting model. We first unify and represent multi-source data by extracting service invocation graphs, and model dependencies between microservices by using a message-passing based graph neural network to generate graph-level sequences. The graph-level representation is then aligned with the LLM, and the LLM is fine-tuned to capture complex spatio-temporal patterns, generating a global vector that represents the state of microservice system within a timeslot. LLM4MST achieves accurate anomaly detection and root cause localization by jointly training the end-to-end model. Experiments on real datasets show that LLM4MST exhibits excellent performance in both full-sample and few-shot scenarios, demonstrating the powerful ability of LLMs in cross-domain knowledge transfer and few-shot learning.