HyperGen: Optimizing Generative Inference with Long Prompts for Resource-Constrained Systems
Lingwen Gong, Kaixin Liu, Xiaolu Li, Shujie Han, Patrick P. C. Lee, Yuchong Hu, Dan Feng · 2025
Generative inference with long prompts often exceeds GPU memory limits in resource-constrained systems (e.g., workstations and edge devices), thereby causing inference failures. We present HyperGen, a lightweight generative inference framework that optimizes the prefill stage (i.e., when input prompts are processed) via two fine-grained partitioning techniques: (i) partitioning and loading model parameters with size awareness into GPU memory; and (ii) partitioning computations of different inference steps to fit into GPU memory and offloading concatenation of partial results to CPU memory. Evaluation shows that HyperGen supports a maximum prompt length of up to 3.8× longer than an existing GPU-based inference approach and reduces the time-to-first-token from hours to seconds compared to the CPU-based prefill approach.