GLITCHES: GPU-FPGA LLM Inference Through a Collaborative Heterogeneous System
Fan Yang, Xinhao Yang, Hongyi Wang, Zehao Wang, Zhenhua Zhu, Shulin Zeng, Yu Wang · 2024
Large language models (LLMs) demonstrate strong capabilities across various tasks. However, in latency-sensitive scenarios, a small batch or even one batch is usually required. This leads to the prefill and the decode stage of LLM inference being computational and memory bottlenecks, respectively. Therefore, it is difficult for a homogeneous FPGA or GPU system to simultaneously address different computational bottlenecks in different stages of LLM inference, resulting in long prefill latency on FPGAs and low utilization during the decode stage on GPUs. This paper proposes GLITCHES, GPU-FPGA LLM inference through a collaborative heterogeneous system. In this paper, we analyze the different characteristics of GPUs and FPGAs and employ GPUs for the prefill stage and FPGAs for the decode stage, leveraging the strengths of GPUs and FPGAs. Based on HBM profiling results, we apply the data prefetching technique to further improve the off-chip memory bandwidth utilization during the decode computations on FPGAs. Experiments demonstrate that a GLITCHES heterogeneous LLM inference system with an A100 GPU and seven U280 FPGAs achieves a 1.28/1.34 times improvement in system throughput and a 2.38/1.90 times improvement in cost efficiency compared to a homogeneous system with 8-card A100/V100S GPUs.