Smart Anomaly Detection: Deep Learning modeling Approach and System Utilization Analysis
Mourad Bouache, Benaoumeur Senouci · 2022
The objective of this project is to perform automated classification of anomalies of system within a large scale cloud production environment using neural network based models on time series data. In large clusters of mixed workload servers, the ability to automatically identify abnormal system utilization is a challenge due to the scale of the problem. To solve this problem, we use deep learning modeling techniques, Long-Short Term Memory (LSTM) models. The data set, to build and test the models, is from production systems over the span of two weeks. The models will use utilization metrics such as CPU, Memory, Network IO, Process Run Queue and Open Files. Anomalous usage in the production cluster is classified as (1) very low usageless than 5% across selected metrics and (2) known anomalous behaviors like memory leaks. This paper will explain how we can create a model that will identify the anomalies we want to flag, in the real world data. We are using Intel Optimized TensorFlow in containers distributed within a cluster of TensorFlow servers.