BitWeaver: Read-Time Truncation in Memory

Garrett Gagnon, Srikanth Malla, Yangwook Kang, Liu Liu · 2025

Large language models (LLMs) have demonstrated remarkable capabilities in generating contextually relevant responses to prompts, but their inference performance is often constrained by a severe memory bottleneck in the self-attention stage.This bottleneck, which is inherently memory-bound, has led to extensive research into strategies for reducing the size of the Key-Value (KV) cache.Many existing approaches employ quantization to lower the data precision and reduce data volume.However, these methods are constrained by memory technology, which requires the data read from memory to match the data written into it.As a result, such strategies must apply reductions at write-time, limiting either model performance or achievable speedup.We identify that enabling precision scaling at read-time -after data has been stored in memory -offers a unique opportunity to simultaneously reduce memory traffic and retain model accuracy.To this end, we develop a read-time precision-scaling mechanism and introduce BitWeaver, a hardware-enabled solution for in-memory truncation.BitWeaver dynamically reduces data precision during memory reads, achieving up to a 3× increase in memory throughput and execution speedups of up to 80% compared to baseline.Additionally, BitWeaver enhances sparse KV cache strategies by improving the efficiency of state-of-the-art sparsity techniques.

Read the paper · More papers on PaperTik