LLM-CIM: A 28nm 126.7TOPS/W Input-LUT-Based Digital CIM Macro with Reconfigurable Matrix Multiplication and Nonlinear Operation Modes for LLMs
Yiqi Wang, Zhen He, Zihan Wu, Ruiqi Guo, Longke Yan, Huiming Han, Yang Wang, Shaojun Wei, Yang Hu, Fengbin Tu, Shouyi Yin · 2025
This paper presents a Digital Computing-in-Memory (DCIM) macro tailored for Large Language Model (LLM) acceleration, named LLM-CIM. It has three key features: 1) A Maximum Reuse-Oriented Input-LUT (MRIL) CIM reduces compute logic area and power by 40.7 % and 37.9 %. 2) An Outlier Aggregation-based Reordering Unit (OARU) saves$1.68 \sim 1.72\mathrm{x}$CIM computation time. 3) A Normalized-Domain Computation Converter (NDCC) improves resource utilization to$95.6 \sim 98.7 \%$. LLM-CIM achieves 126.7TOPS/W energy efficiency on LLaMA2-7B, with up to$5.36 x$energy saving over SOTA CIMs.