Multi-Tier Workload Consolidations in the Cloud: Profiling, Modeling and Optimization
Kejiang Ye, Haiying Shen, Yang Wang, Chengzhong Xu · IEEE Transactions on Cloud Computing · 2020
Reducing tail latency becomes increasingly important to improve the user-perceived service experience. User-facing latency-sensitive cloud applications typically contain multiple interactive tiers (e.g., Web, App, Database) running in different virtual machines (VMs) with complex interaction patterns. However, such interactions between VMs in different tiers are often neglected in previous VM consolidation methods, resulting in poor application performance. In this article, we study the consolidation of multi-tier interactive workloads from a new perspective of user-perceived tail latency. We propose a novel profiling-based consolidation methodology to satisfy tail latency requirements while reducing the number of used physical machines. To achieve such a goal, we first perform large-scale profiling experiments under various consolidation settings in a KVM virtualized private cluster to establish the empirical performance values. We consider two key factors that affect the tail latency of multi-tier workloads:interferencewith co-located VMs andinteractionbetween tiers. We model the consolidation of multi-tier workloads as an optimization problem with different objectives and constraints, and derive the consolidation schedule. We implement and evaluate the proposed models, as well as comparing with other methods (i.e.,withoutprofiling orwithoutconsidering interaction influence). Extensive experimental results show that the proposed method is able to reduce up to5Xtail latency, compared with the methodwithoutprofiling and up to1.3Xtail latency, compared with the methodwithoutconsidering the interaction influence between different tiers.