Quartet: A Digital Compute-in-Memory Versatile AI Accelerator With Heterogeneous Tensor Engines and Off-Chip-Less Dataflow
Yikan Qiu, Guoxiang Li, Meng Wu, Yifan Jia, Le Ye, Yufei Ma · IEEE Transactions on Circuits and Systems I Regular Papers · 2025
Although most AI core operations can be formulated as matrix multiplications (MMs), their characteristics are quite different. In some cases, one input may be a constant weight matrix while the other is a dynamic feature matrix, i.e., FWMM, or both inputs may be dynamic features as FFMM. Furthermore, the data involved can also be sparse, leading to variations like SpFWMM and SpFFMM. To address this challenge, this paper investigates a versatile accelerator architecture for AI algorithms based on heterogeneous tensor engines. For MM operators with varying characteristics, this paper proposes four-quadrant heterogeneous tensor engines to handle FWMM, SpFWMM, FFMM, and SpFFMM, respectively. These four tensor engines are comprised of SRAM-based single address digital compute-in-memory (CIM) array, SRAM-based multi-address digital CIM array, systolic array, and multi-SIMD array, respectively. In addition, to improve the AI execution efficiency, this paper proposes a dual-level multi-issue mechanism to achieve inter-operator and inter-block parallelization along with an off-chip-less dataflow enabled by on-chip unified memory pool. Thanks to the integration of aforementioned innovations, this paper develops a versatile AI acceleration chip Quartet, which achieves exceptional energy efficiency. Specifically, for graph convolutional network on PubMed, it demonstrates a$19.56\times $and$3.47\times $improvement in energy efficiency compared to similar works, ReDCIM and TensorCIM, respectively.