Exploring Approaches to Integrate Performance Prediction and Anomaly Detection in Microservices Systems
Hamza Hussain, Ghadeer Abuoda, Marin Litoiu · 2024
Numerous algorithms have been proposed over the years to predict performance metrics and detect performance anomalies in microservices based applications. However, most models specialize in either performance prediction or anomaly detection. As a result, multiple models are often required to monitor cloud-native applications effectively. Given the distributed nature of modern cloud-native systems, performance and health monitoring is carried out through various channels, and using separate models for each task adds complexity to the monitoring process. Therefore, this paper aims to investigate the use of multimodal data to integrate performance prediction and anomaly detection methods. For this purpose, we utilize simple Graph Neural Networks to predict latency distribution, as opposed to a single latency value for traces generated by a microservices system. The predicted latency distribution is then fed to several Machine Learning models to predict trace based anomalies. Our results show that all models tested can achieve accuracy of more than 96% and with good precision and recall values in a heavily unbalanced Train Ticket dataset.