A 2.9–33.0 TOPS/W Reconfigurable 1-D/2-D Compute-Near-Memory Inference Accelerator in 10-nm FinFET CMOS
H. Ekin Sumbul, Gregory K. Chen, Phil C. Knag, Raghavan Kumar, Mark A Anders, Himanshu Kaul, Steven Hsu, Amit Agarwal, Monodeep Kar, Seongjong Kim, Ram Kumar Krishnamurthy · IEEE Solid-State Circuits Letters · 2020
A 10-nm compute-near-memory (CNM) accelerator augments SRAM with multiply accumulate (MAC) units to reduce interconnect energy and achieve 2.9 8b-TOPS/W for matrix-vector computation. The CNM provides high memory bandwidth by accessing SRAM subarrays to enable low-latency, real-time inference in fully connected and recurrent neural networks with small mini-batch sizes. For workloads with greater arithmetic intensity, such as large-batch convolutional neural networks, the CNM reconfigures into a 2-D systolic array to amortize memory access energy over a greater number of computations. Variable-precision 8b/4b/2b/1b MACs increase throughput by up to 8× for binary operations at 33.0 1b-TOPS/W.