WS-CIM: Enabling Fast and Simultaneous Update for Multi-Macro Compute-in-Memory Architecture Using Weight Sharing Technique
Yan-Ding Shieh, Ming-Guang Lin, Hung‐Yu Wang, An-Yeu Andy Wu · 2025
Compute-in-memory CIM) architecture has been widely studied to accelerate deep neural networks (DNNs). CIM improves energy and area efficiency by performing multiply-accumulate (MAC) operations within memory array. However, the limited size of individual CIM macros requires frequent weight updates for large DNN models, leading to significant latency and energy overhead. In this paper, we propose WS-CIM, a novel framework to enable fast and simultaneous weight updates for multi-macro CIM architectures. Specifically, WS-CIM adopts fine-grained weight sharing technique considering the CIM architecture to minimize redundant write operations, guided by a distance loss function to maintain accuracy. A workload balance mechanism is further introduced for multi-macro architecture to prevent bottlenecks from any single macro during simultaneous weight updates. Experiments on ResNet50 and DeiT-B show that WS-CIM achieves 34.8% latency reduction with 1.29% accuracy loss for ResNet50 and 35.8% latency reduction with 0.70% accuracy degradation for DeiT-B, respectively. These results demonstrate the scalability and efficiency of WS-CIM for deploying DNNs on advanced CIM hardware platforms.