Toward Autonomic Cloud: Automatic Anomaly Detection and Resolution
Rafiul Ahad, Eric S.W. Chan, Adriano dos Santos · 2015
In this paper we describe an approach to implement an autonomic cloud. Our approach is based on our belief that if a computing system can automatically detect and correct anomalies - including response time anomalies, load anomalies, resource usage anomalies, and outages - then it can go a long way in reducing human involvement in keeping the system up, and that can lead to an autonomic system. We focus on a class of anomalies that are defined by normal values expected of key metrics. We describe a hierarchical rule-based anomaly detection and resolution framework for such a class of metrics.