SyncIntellects: Orchestrating LLM Inference with Progressive Prediction and QoS-Friendly Control

Xue Lin, Zhibo Zhang, Peining Yue, Haoran Li, Jin Zhang, Baoyu Fan, Huayou Su, Xiaoli Gong · 2024

Large Language Models (LLMs) have shown impressive capabilities, especially in the realm of Human-Machine Chat Systems. Nevertheless, these models entail significant computational expenses, particularly when generating tokens. As a remedy to enhance system throughput and hardware utilization, batch scheduling is commonly adopted. This method involves initiating a batch of inference requests concurrently and then waiting for their completion. A significant challenge encountered with task-batching is the need to group requests with similar response lengths. However, accurately predicting response length proves to be a daunting task, and the inherent variability in response length leads to suboptimal resource utilization.In this paper, we introduce SyncIntellects, a framework designed to orchestrate Large Language Model (LLM) Inference with fine-grained response length prediction and Quality of Service (QoS)-Friendly length control. Specifically, SyncIntellects enhances response length prediction by leveraging embedding information during token generation through a transformer-based model. Subsequently, a dynamic response length controller based on Prompt Engineering techniques is employed to ensure alignment of response lengths without compromising the QoS of the responses. We have implemented SyncIntellects and seamlessly integrated it with a chatbot engine based on the llama2 7B model. We conduct comprehensive experiments on an NVIDIA A100-based testbed, and the results demonstrate a significant reduction in latency by 17.76% on average, along with an increase in throughput by 9.34%.

Read the paper · More papers on PaperTik