QUART: Latency-Aware FaaS System for Pipelining Large Model Inference
Yanying Lin, Yanbo Li, Shijie Peng, Yingfei Tang, Shutian Luo, Haiying Shen, Chengzhong Xu, Kejiang Ye · 2024
Pipeline parallelism is a key mechanism to ensure the performance of large model serving systems. These systems need to deal with unpredictable online workloads with low latency and high good put. However, due to the specific characteristics of large models and resource constraints in pipeline parallelism, existing systems struggle to balance resource allocation across pipeline stages. The primary challenge resides in the differential distribution of requests across various stages of the pipeline. We propose QUART, a large model serving system that focuses on optimizing the performance of key stages in pipeline parallelism. QUART dynamically identifies the key stages of the pipeline and introduces an innovative two-level model parameter caching system based on forks to achieve rapid scaling of key stages within seconds. In evaluations with real-world request workloads, QUART reduces average response latency by up to 87.1%) and increases good put by 2.37x compared to the baseline. The experiments demonstrate that QUART effectively reduces tail latency and the average queue length of the pipeline.