Universal Anomaly Detection Method Based on Massive Monitoring Indicators of Cloud Platform

Min Li, Dingyong Tang, Zepeng Wen, Yunchang Cheng · 2021

The cloud platform concept has been regarded as the mainstream architecture of large-scale service systems. It can support different types of services and improve the stability and scalability of serving. Since the cloud platform is composed of a large number of basic services, the quality of its maintenance has gradually become the research focus of daily work. Performance monitoring indicators can completely describe the current status of the service. But due to the huge number of them, it is often impossible to identify the real-time root causes only by empirical alarm threshold. Besides, when a failure breaks out, it usually shows abnormality on different indicators, which brings great difficulty to artificially searching for the real root causes. This paper introduces a universal anomaly detection algorithm, which is based on the performance monitoring data of two different real cloud platform systems. We used both offline and online learning methods to dynamically revise the feature matrix and anomaly threshold of indicators. The feature matrix will be divided into time periods in order to improve the accuracy. Then, we used deviation degree, scoring and sorting algorithm to identify the root causes of faults in real-time data. The model was verified on the dataset of the AIOps Challenge (2021), which is an international operation and maintenance competition. During the experiment, we discovered a hyperparameter that greatly affects the quality of detection, then we tested it for times to find a better result. Finally, it was proved that our universal model has a good performance on the dataset from two different application systems.

Read the paper · More papers on PaperTik