A 29.12-TOPS/W Vector Systolic Accelerator With NAS-Optimized DNNs in 28-nm CMOS

Kai Li, Mingqiang Huang, Ang Li, Shuxin Yang, Quan Cheng, Hao Yu · IEEE Journal of Solid-State Circuits · 2025

The increasing model size and computational load of deep neural networks (DNNs) present a significant challenge to deploy DNN models on constrained devices. To enhance performance without compromising accuracy, we introduce a neural architecture search (NAS) method to develop layer-wise mixed-precision and mixed-sparsity DNNs. However, the optimization cannot be directly applied to existing DNN accelerators due to the specific data requirements of layer-wise mixed precision and mixed sparsity. To address this issue, this research proposes separate mixed-precision and mixed-sparsity accelerators. Both accelerators demonstrate cutting-edge results. The mixed-precision accelerator utilizes a split-and-combination vector (SCV) to re-use variable precision units (1/2/4/8 bit) at each layer with a further vector systolic array (VSA) implementation. By optimizing NAS for mixed-precision VGG-16, the VSA achieves mixed energy efficiency reached 29.12 TOPS/W, which is equivalent to 2 bit and equivalent accuracy at 4 bit. The mixed-sparsity accelerator introduces a log-scale structured sparse encoding strategy, combined with MAC and group VSA (G-VSA) optimization, to enhance system performance. It achieves an average energy efficiency of up to 21.7 TOPS/W at 0.7 V and 400 MHz using 28-nm CMOS. The measured results show that the mixed-precision chip exhibits better energy efficiency and accuracy than the mixed-sparsity chip.

Read the paper · More papers on PaperTik