Power-Controlled Job Slowdown in Data Centers
Ariana J. Mann, Nicholas Bambos · 2022
As compute demands soar upwards, it is essential that systems not only scale throughput and latency, but also improve energy efficiency and alleviate power bottlenecks. A common quality of service metric for user-facing data center services is slowdown, the ratio of the actual response time of a computational job to the expected response time under no delay. In this work, we expand the definition of slowdown to the power-controlled, rate-optimizing setting. We take the initial steps to investigate and analyze the key tradeoffs between power and slowdown costs. Formulating the problem in a dynamic programming framework, we provide analytic solutions in the single class job setting, and formulate the multi-class job setting. Numerical solutions to the dynamic programming equation are also produced, and confirm the analytical results and parameter tradeoffs. In particular, we demonstrate both analytically and numerically that reasonable tradeoffs can be made between power and slowdown by optimizing the server processing rate. Optimizing for power consumption is increasingly important as data centers’ tackle the transition to renewable energy sources.