Activation Sequence Caching: High-Throughput and Memory-Efficient Generative Inference with a Single GPU

Sowoong Kim, Eunyeong Sim, Youngsam Shin, Yeongon Cho, Woongki Baek · 2024

Generative artificial intelligence is widely used for various tasks such as language translation and art creation. Generative inference employs the key and value tensors to encode relational information among the tokens in the input and output sequences. Most of the existing generative inference frameworks use KV caching (KVC), which caches the key and value tensors (i.e., the KV cache) to avoid recomputing the key and value tensors for each of the processed tokens. Despite the widespread use of KVC, in-depth characterization of KVC with tensor offloading, which enables generative inference with a single GPU, remains yet to be explored.

Read the paper · More papers on PaperTik