Resource profiling for large-scale data centres

Christopher B. Hauser · OPen Access Repositorium der Universität Ulm (OPARU) (Ulm University) · 2021

The use of virtualisation allows to share physical resources in a data centre among multiple tenants in parallel. Cloud Computing became the de facto standard in modern data centre management. Yet, resource interferences and non-ideal placement decisions can hinder an equally distributed and fair shared data centre utilisation for the participants, the data centre provider and the customers. Related work is improving virtualisation and isolation, works on resource management in data centres, addresses green computing aspects to improve the ecological aspects of data centres, reviews and improves data centre monitoring, and works on time series analysis. The popular data centre management modes are Cloud Computing or High Performance Computing. Both have in common to consist of a centralised management component, which schedules and allocates resources, based on static criteria. A dynamic resource allocation is more complex and expensive, since utilisation-aware scheduling requires monitoring and processing to express resource demands as profiles to resource management components. This thesis presents solutions for the monitoring and profiling, to build a distributed dynamic resource allocation for large-scale data centres. The designed and implemented distributed monitoring for virtualised nodes in shared data centres works cross-layer while being non-intrusive, elastic and robust. The DisResc Monitoring proposed here considers static and dynamic metrics of physical and virtual layer with multi-tenancy awareness, and introduces a flexible and scalable communication model. The DisResc communication uses a distributed message bus with publish-subscribe, to allow fully distributed, hierarchical, or centralised setups. The kvmtop collector, as proof of concept implementation, transmits comprehensive static and dynamic metrics for virtual machines on KVM hypervisors, including runtime overhead and with awareness of overbooking. The evaluation shows that this black box monitoring approach produces accurate measurements, equivalent to monitoring inside the virtual machines. Resource interferences, due to overbooking and high utilisation values, can be detected accurately by the monitoring design and by kvmtop. The designed and implemented distributed profiling selects statistical methods and probability theory to process the monitoring data as distributed profiler instances, next to the monitoring collector on each node in the data centre. The DisResc Profiler therefore requests and subscribes to events from the DisResc Monitoring instances, to sequentially process the online time series stream. The utilisation values are aligned to the static hardware properties, discretised to states, and transformed to Markov chains. Transition matrices of the Markov chain are built for the overall profile, for each period in a period tree with fix size, and for each of automatically detected phases. The period tree dimension is configured using pre-processing steps like signal processing methods. The phase detection uses the recorded probabilities to differentiate patterns by likeliness of a sequence of states. The prototypical implementation TSProfiler provides tools to produce profiles, simulate utilisation values from profiles, and compute the likeliness of future utilisation values. The TSProfiler tools are used to evaluate the profiling approach. An n-step-ahead prediction on existing data sets calculates the error of predicted and actual utilisation values. The prediction error is significantly lower than using an overall average. The accuracy and the computation complexity depends on the profiling parameters, like the number of states, which mode is used (overall transitions, with period tree, with phases), and the number of prediction steps. The profile is distributed via the DisResc communication, and contains hardware-independent representations of CPU, disk and network utilisation. The presented DisResc monitoring and profiling provide necessary information as input for a dynamic resource allocation. The distributed, non-intrusive approach for monitoring virtual environments and processing the stream directly allows to deploy the components as independent instances decentralised in large-scale data centres. The profiling compresses the raw monitoring data, and produces expressive representations, suitable for direct use to derive resource allocation decisions. Monitoring and profiling are designed for low resource demanding, and the resulting profile is designed to be lightweight. Compared to state of the art centralised monitoring and analysis tools, the presented thesis approach follows a decentralised approach to achieve horizontal scalability.

Read the paper · More papers on PaperTik