Optimizing application downtime through intelligent VM placement and migration in cloud data centers
Venkatesh Nandakumar, Alan Wen Jun Lu, Madalin Mihailescu, Zartab Jamil, Cristiana Amza, Harsh Vikram Singh · Computer Science and Software Engineering · 2015
As cloud data centres grow in size and complexity, hosted applications become increasingly vulnerable to dynamically occurring infrastructure downtime periods caused by partial infrastructure failures. Downtimes within cloud data centres can be diverse, ranging from unplanned server/rack unit failures to compulsory server power-offs when addressing arbitrary environment conditions, e.g., thermal issues. For instance, in these environments, server racks are often a unit of failure due to either faulty rack switches or rack power units. We observe that the degree of application disruption depends on i) the application's fault tolerance, reconfiguration capabilities, and redundancy of VM components affected by the respective emergency shutdowns and ii) the support for VM migration of vulnerable application components within the constrained time window of impending shutdown of a failure unit. In this context, in this paper, we develop and evaluate techniques which aim to optimize the downtime of hosted applications during emergency shutdowns due to partial failures through two orthogonal approaches: i) designing VM placement techniques that are aware of application fail-over semantics and ii) prototyping intelligent schemes for live VM migration prioritization based on prediction models for both VM migration times and expected downtimes for applications.