CLAT: A Clustering-Based Attention Transformer Accelerator for Low-Latency Text Generation in LLMs

Sunwoo Lee, Beomseok Kim, Jeongwoo Park, Dongsuk Jeon · IEEE Transactions on Circuits and Systems I Regular Papers · 2025

Transformer-based large language models (LLMs) excel in text generation but face challenges like memory bandwidth bottlenecks and large key-value (KV) cache sizes as context lengths grow, impacting low-latency performance. Existing accelerators adopt parallelism, model compression, and sparsity exploitation but often fail to fully utilize token-specific sparsity, limiting their effectiveness for long context lengths. CLAT addresses these issues with a low-overhead clustering algorithm that identifies relevant key vector clusters for each token’s query, omitting less relevant vectors with minimal impact. It optimizes memory bandwidth using routing for single-batch inference and introduces scheduling techniques to reduce attention layer latency. Additionally, CLAT compresses model parameters to 4-bit precision and KV caches to 8-bit precision, supported by a multi-precision MAC structure that avoids extra overhead. Validated on Llama2-7B, OPT-6.7B, and Llama3-8B models, CLAT reduces attention layer latency by up to 88.6% and overall text generation latency by up to 34.9%. It improves single-batch text generation throughput by$1.66\times $to$2.42\times $over an A100 GPU, demonstrating significant performance gains.

Read the paper · More papers on PaperTik