PS-IMC: A 2385.7-TOPS/W/b Precision Scalable In-Memory Computing Macro With Bit-Parallel Inputs and Decomposable Weights for DNNs
Amitesh Sridharan, Jyotishman Saikia, Anupreetham Anupreetham, Fan Zhang, Jae-sun Seo, Deliang Fan · IEEE Solid-State Circuits Letters · 2024
We present a fully digital multiply and accumulate (MAC) in-memory computing (IMC) macro demonstrating one of the fastest flexible precision integer based MACs to date. The design boasts a new bit-parallel architecture enabled by a 10T bit-cell capable of four AND operations and a decomposed precision data-flow that decreases the number of shift-accumulate operations, bringing down the overall adder hardware cost by 1.57x whilst maintaining 100% utilization for all supported precision. It also employs a carry save adder tree that saves 21% of adder hardware. The 28nm prototype chip achieves a speed-up of 2.6×, 10.8×, 2.42×, and 3.22× over prior SoTA in 1bW:1bI, 1bW:4bI, 4bW:4bI, and 8bW:8bI MACs respectively.