Configurable Dataflow and Adaptive Mapping Optimization for Hybrid ReRAM and SRAM Compute-in-Memory Accelerator
Jingyu He, X. S. Wang, Kunming Shao, Kwang-Ting Tim Cheng, Chi-Ying Tsui · IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 2025
Hybrid compute-in-memory (CIM) designs have been proposed recently to facilitate the storing of large number of weights of a neural network on-chip. Notably, ReSCIM wang2024res pairs an SRAM cell with a dedicated ReRAM crossbar, allowing ReRAM to serve as the local storage, significantly enhancing the storage capacity of the SRAM-CIM. The SRAM is custom-designed not only to serve as a storage element for CIM but also to function as a sense amplifier to retrieve the data from the ReRAM, which enables super high bandwidth of weight data loading into the CIM engine. However, existing mapping tools for CIM are inadequate for ReSCIM since they do not fully exploit the unique hardware characteristics and advantages of this novel architecture. In this work, we propose an analytical energy and latency model, which incorporates four key factors: hardware, workload, dataflow, and mapping (HWDM), for executing inference of neural network on the ReSCIM accelerator. Specifically, we first characterize the ReSCIM accelerator hardware specifications and the neural network layers. Next, we introduce three dataflows for ReSCIM, leveraging the high weight-loading bandwidth to reduce memory access for various workloads and layer types. Finally, we develop an algorithm to generate optimal mapping and dataflow strategies aimed at minimizing latency or energy consumption. Using our HWDM model, we design a tile-based ReSCIM accelerator and conduct extensive simulations to obtain the cycle-accurate latency and gate-level energy consumption metrics for inference across different neural networks. We conduct design space exploration (DSE) using the HWDM model on a comprehensive set of benchmarks to minimize inference energy or latency. Experimental results show that our optimal ReSCIM accelerator achieves a 44% reduction in EDP reduction compared to the weight-stationary and fixed mapping baseline for SEResNet50. Moreover, our design exhibits 1.74× higher energy efficiency than the state-of-the-art hybrid TL-nvSRAM wang2023tl accelerator on ResNet 18.