FATHOM: Fast Attention Through Optimizing Memory
Elliott Binder, Arvind Sudarsanam, Ravi Sunkavalli, Tze Meng Low · 2025
Transformer models are built on attention and feedforward layers that are predominantly matrix-matrix multiplication. Although matrix multiplication is often thought to be compute-bound, the matrix dimensions in attention are too small to reach peak compute throughput on many of today's CPU and GPU architectures. These same routines in many dense linear algebra libraries also do not reach the memory bound for matrix multiplication for these sizes. To improve the memory bandwidth utilization of transformer models, we employ a bandwidth-friendly data layout of intermediate data between operations and redesign our matrix multiplication kernels to optimize for memory utilization and bandwidth efficiency. We present Fast Attention Through Optimizing Memory (FATHOM), which achieves up to$6.7\times$higher throughput in batch matrix multiplications and$1.8\times$speedup in end-to-end models on CPU and GPU architectures.