A 68 TOPS/W, 256MB SRAM Sparse GEMM Accelerator Tiled Across 16, 4nm Near Memory Compute (NMC) Chiplets Disaggregated 2.5D System
Srivatsa Srinivasa, Prerna Budhkar, Gauthaman Murali, Vui Seng Chua, Paolo A. Aseron, Vinayak Honkote, Ravishankar Iyer, Nilesh K. Jain, Dileep John Kurian, Anuradha Srinivasan, Tanay Karnik · 2025
The rapid evolution of AI models from DNNs to Transformers, characterized by their increasing size and complexity, presents a significant challenge for hardware acceleration. The core computational workload Matrix Multiplication (MatMul) demands immense computational power, while the burgeoning model parameters and sequence lengths necessitate substantial memory capacity [1]. Traditional hardware architectures struggle to meet these demands, leading to performance bottlenecks and energy inefficiency [2]. This requires a holistic design approach, encompassing specialized hardware accelerators, optimized memory hierarchies, and novel data movement techniques. Leveraging chiplet disaggregation to push the boundaries of hardware technology co-design will overcome the inefficiencies of monolithic hardware accelerators and drive transformative innovations. We present 20 chiplet disaggregated Sparse GEMM (SPGEMM) acceleration solution with 16 Near Memory Compress Compute (NMC) chiplets offering a total of 68 TOPS/Wand 256MB SRAM minimizing frequent off package data movement.