An Efficient DNN Model Serving System using Layer-wise Caching and Direct-Host-Access

Jinwoo Jeong, Jeongseob Ahn · ACM Transactions on Computer Systems · 2025

With the increasing demand to utilize deep neural networks (DNNs) in online services, it is important to serve DNN models on GPUs in a cost-effective manner. Once the required DNN model is ready in the GPU memory, we can immediately serve the inference requests with low latency. Otherwise, it needs to load the model from host to GPU, adding a significant delay to inference. This article proposes Ignite to minimize cold-start latency while provisioning DL models from host to GPU in server environments. First, we propose LCache to effectively utilize the limited GPU memory for model serving with unique cache replacement policies. We devise layer-wise cache replacement policies tailored for executing DL inferences in a pipelined way. Second, we take advantage of the direct-host-access facility provided by commodity GPUs, allowing access to particular layers of models in the host memory directly from GPU without loading. We show that Ignite can effectively reduce the cold-start latency while increasing the throughput of serving DNN models. When deploying multiple ResNet, BERT, and RoBERTa instances on a DL inference serving system, Ignite shows a significant performance improvement compared to the pipelining technique and stable 99% tail latency.

Read the paper · More papers on PaperTik