A 28-nm RRAM/SRAM Collaborative CIM Accelerator Supporting RRAM-Endurance-Latency Awareness for Edge Fine-Tuning

Chen Mu, Zhirui Huang, H. W. Jiang, Jie Liao, Yuliang Leon Zhou, Liang Chen, Yechu Zhang, Haozhe Zhu, Jianguo Yang, Qi Liu, Chixiao Chen · IEEE Journal of Solid-State Circuits · 2025

The resistive random access memory (RRAM)-based computing-in-memory features high density and high energy efficiency on edge. However, fine-tuning RRAM-based SoCs remains challenging due to the inherent limitations of non-volatile memory (NVM) characteristics. This work proposes an NVM-endurance/latency-aware collaborative RRAM/static random access memory (SRAM) compute-in-memory (CIM) accelerator that addresses the difficulties associated with RRAM write operations during fine-tuning tasks for both CNN and transformer models. The major contributions are: 1) RRAM-most significant bit (MSB)–SRAM-least significant bit (LSB) (RMSL)-based collaborative CIM macros, mitigating RRAM cell flipping times and alleviating the endurance concern; 2) an RRAM-sparse-SRAM-dense (RSSD) weight updating engine, minimizing the long reading and writing latency associated with RRAM access; and 3) a row-wise pipeline weight gradient (WG) computing data flow with low-hardware overhead. With a maximum bit update of ten times per fine-tuning (20 epochs, from start to convergence), the system achieves an energy efficiency of 76.25 TOPS/W for the CIM macro and 22.07 TOPS/W for the overall system. Thanks to the bit-level collaborative CIM, the RRAM CIM macro achieves the same hardware utilization during fine-tuning and inference processes to support RRAM’s energy-efficient computing. The proposed CIM accelerator, fabricated using 28-nm CMOS technology, achieves up to$143{\times }$RRAM endurance improvement,$117{\times }$and$144{\times }$reduction in RRAM write power and latency.

Read the paper · More papers on PaperTik