Foreseer: Knowledge-Driven Acceleration of Memory-Bound Matrix Multiplications for Large Language Model Inference

Cong Li, Yutao Xu · 2024

The majority of the latency in large language model inference lies in a few memory-bound matrix multiplication kernels. Traditionally, those kernels are optimized by extensive offline experiments of different tiling, striding, and slicing parameter choices on the target hardware. Scaling to different models, inference settings, and GPU hardware then becomes difficult. In this paper, we propose Foreseer, a new policy to generate the high-performing parameters for different kernels on the fly without offline experiments on the hardware. Foreseer leverages the specific characteristics of memory-bound matrix multiplication kernels, and synthesizes various prior knowledge such as keeping the adequate memory access pressure, avoiding the potential tail effect from load imbalance in dispatching, etc., into a simple penalty function. For a given parameter choice, Foreseer provides an analytical estimation of the impacting factors such as the memory access pressure, the potential tail effect, etc., to score the choice. A high-performing parameter choice is then discriminated from the penalty function. Micro-benchmarks of matrix multiplication kernels from popular language models on two different high-end Intel Data Center GPUs demonstrate that Foreseer outperforms different baselines including the vendor libraries, achieving >90% of the best-known performance obtained from the exhaustive offline parameter search experiments. End-to-end inference experiments on those language models with different settings also show that Foreseer accelerates the baseline by 20% on average, being competitive with and quite often exceeding the performance of the comprehensively-optimized vendor inference engine.

Read the paper · More papers on PaperTik