Evaluating Compute in Memory Architectures for Matrix Multiplication: A Dataflow-Centric Perspective
Tanvi Sharma, Indranil Chakraborty, Mustafa Ali, Kaushik Roy · 2025
Compute in memory (CIM) is a promising technique to reduce data movement costs in traditional hardware by efficiently performing in-situ matrix multiplication, the dominant computation during deep learning (DL) inference. However, the broader question of how CIM architectures compare to tensor-core-like architectures remains largely unexplored. In this work, we take a dataflow-centric approach utilizing classic parameters such as compute latency, bandwidth, capacity and compute/memory access costs to determine throughput and energy consumption of a given CIM architecture. To that effect, we perform an iso-area comparison of tensorcore-like (or PE array) architecture with different CIM integrated architectures for matrix multiplication kernels. Our results demonstrate that CIM integrated memory can improve energy efficiency by up to$3.5 \times$and throughput by up to$11 \times$compared to tensorcore baseline, considering INT8 precision.