Large scale anomaly detection in data center logs and metrics
Rafael P. Martínez-Álvarez, Carlos Giraldo-Rodríguez, David Chaves-Diéguez · 2018
Data centers continuously produce large amounts of data related to their internal operation. This kind of machine-generated data is flowing 24x7x365; however, it is seldom exploited to benefit the health of the processes and the business itself. The information usually comes in two flavors: application events or system logs, and periodic measurements of some changing magnitude (processor load, used memory, etc.). We have, therefore, a mix of structured and unstructured data with a high intrinsic variety, yet containing a high strategic value for those able to extract it. In this work, we propose a data processing engine for anomaly detection based on a real setup made for an IT services company who wanted to enhance its portfolio of technological solutions. Two were the main challenges we faced as fundamental requirements: making the system work in tough big data environments, and being able to yield accurate real-time responses.