Deferred Continuous Batching in Resource-Efficient Large Language Model Serving
Yongjun He, Yao Lu, Gustavo Alonso · 2024
Despite that prior work of batched inference and parameter-efficient fine-tuning techniques have reduced the resource requirements of large language models (LLMs), challenges remain in resource-constrained environments such as on-premise infrastructures to serve workload that is composed of both inference and fine-tuning jobs. Prior solutions must either pause existing jobs which causes service interruptions, or queue new jobs which results in a long delay.