CoFRIS: Coordinated Frequency and Resource Scaling for GPU Inference Servers
Marcus Chow, Daniel Lin-Kit Wong · 2023
Data centers have a variety of metrics that they must adhere to. Not only do they have to meet the rate of incoming requests, but each request also has a service level objective (SLO) that they must satisfy. However, the average latency of a single request typically is much faster than the tail latency of the SLO. This creates a latency slack gap between the average and tail latencies. This latency slack can be exploited to reduce power by slowing down requests through a variety of techniques, such as frequency and resource scaling. However, we show that in an inference server context, frequency alone cannot slow down a request far enough, leaving slack left to be explored. To make up this slack, we propose CoFRIS, a coordinated frequency and resource scaling effort for GPU inference servers. CoFRISdynamically configures the GPU frequency and active resources to minimize power while meeting variable throughput and latency demands. We evaluate CoFRISwith compute unit (CU) level power gating and improve power consumption by 28% over no frequency or resource scaling, 13% improvement over using only frequency scaling, and 5% over using only CU resource scaling.