Reconfigurable Precision INT4-8/FP8 Digital Compute-in-Memory Macro for AI Acceleration
Jinane Bazzi, Mohammed E. Fouda, Ahmed M. Eltawil · 2025
Compute-in-memory (CIM) technology has emerged as a promising solution to address the computational demands of deep neural network (DNN) models, which require substantial multiply-accumulate (MAC) operations. However, there is a growing need for reconfigurable CIM architectures that can support both integer (INT) and floating-point (FP) operations within a single design. This flexibility is crucial for optimizing efficiency and resource utilization, especially in DNN applications involving mixed-precision computations. In this work, we propose a reconfigurable precision digital macro design that accelerates MAC computations while supporting INT4, INT8, and FP8 configurations within the same architecture. Both signed and unsigned operations are supported in INT mode. To enhance performance, the proposed design uses a parallel-input approach and a mantissa parallel-alignment technique in FP mode. The macro is implemented in 40nm CMOS technology. It achieves a peak throughput of 7123.48 GOPS in INT4 mode and 1187.25 GFLOPS in FP8 mode, with peak energy efficiencies of 367.45 TOPS/W and 23.14 TFLOPS/W, respectively.