AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUs
Shiheng Cao, Junmin Wu, Junshi Chen, Hong An, Zhibin Yu · 2025
Large language model (LLM) inference applications are surging in recent years, which largely relies on modern GPUs.On the other hand, GPU analytical model is a commonly used tool for architects to precisely identify bottlenecks quickly with deep insights.However, existing GPU analytical models fall short of accurately modeling LLM inference applications on modern GPUs, because of unsuitable tensor core modeling, ignoring constant cache as well as instruction cache modeling and abstracting away important details for LLM inference applications.To address this problem, we propose a novel analytical model dubbed AMALI to accurately model LLM inference on modern GPUs with three innovations.First, we develop an instruction modifier and throughput based tensor core model by accurately capturing the math pipe throttle stalls to enhance the architecture modeling for modern GPUs.Second, we propose analytical models for constant cache and instruction cache by developing micro-benchmarks to measure CUDA kernel launching latencies.This significantly improves AMALI's accuracy compared to real GPU hardware.Finally, we design a multi-warp model by leveraging warp instruction number distribution to reflect LLM inference application characteristics.We validate AMALI on an A100 GPU by using typical LLM inference applications.The results show that AMALI reduces the MAPE (mean absolute percentage error) from 127.56% to 23.59%