A 28nm 20.9-137.2 TOPS/W Output-Stationary SRAM Compute-in-Memory Macro Featuring Dynamic Look-ahead Zero Weight Skipping and Runtime Partial Sum Quantization
Xiaofeng Hu, Han-Gyeol Mun, Jian Meng, Yuan Liao, Amitesh Sridharan, Jae-sun Seo · 2025
SRAM-based compute-in-memory (CIM) designs have shown im-pressive energy efficiency for vector-matrix multiplications of AI work-loads [1]–[9]. As many prior CIM works primarily focus on macro-level energy-efficiency, there are often big gaps between macro vs. system energy-efficiencies in real applications [10], [11]. Conventional SRAM CIM macros naturally adopt the weight-stationary (WS) dataflow, and computed partial sums are communicated to SRAM buffers outside of the CIM macro. This largely degrades the WS-CIM system energy-efficiency, especially due to frequent high-precision partial sum (Psum) read/write operations (Fig. 1). Further, while weight sparsity in AI models can be high, CIM macros with WS dataflow make zero weight skipping infeasible due to shared acti-vation and static weight storage. Previous works only support lim-ited structured (e.g. block-wise) sparsity [4], or only store non-zero weights with indices but do not have latency benefits [7]. Several systolic array designs presented skipping techniques for sparse ma-trices [12]–[14], but they all require large registers and complex con-nections, which are unsuitable for compact CIM macros. To ad-dress these challenges, this paper presents the first output-stationary SRAM digital CIM (OS-CIM) macro and system, featuring (1) an OS-CIM structure with custom single-cycle read-and-write 8T SRAM cell to store the Psum in the CIM macro and perform the accumulation locally, eliminating the expensive Psum access to SRAM buffers, (2) a dynamic look-ahead weight skipping (DLS) scheme, well aligned with OS-CIM, which achieves high latency reduction across various weight sparsity patterns and significantly improves workload balance, and (3) a runtime Psum quantization (RPQ) method to mitigate the hardware cost of high-precision accumulation.