9Ring: A 3D-Stacked Memory-Based Accelerator for Flexible and Efficient Deep CNN Applications
Wen Cheng, Qianya Cheng, Yi Liu, Lingfang Zeng, André Brinkmann, Yang Wang · ACM Transactions on Architecture and Code Optimization · 2025
The massive computational and memory requirements of deep convolutional neural networks (DCNNs) have led to the development of neural network (NN) accelerators. However, as DCNN models grow in size, the demands on NN accelerators in terms of performance, memory bandwidth, and power efficiency continue to increase. We, therefore, present 9Ring , a flexible and efficient DCNN accelerator that takes full advantage of 3D-stacked memory, focusing on its hardware architecture, software scheduling, and optimization strategy. In particular, we first show that the mismatch between DCNN accelerators and DCNN models can lead to increased energy consumption and performance bottlenecks. We then present three flexible dataflow scheduling strategies to mitigate this mismatch. Afterward, we introduce an energy efficiency analysis tool that can automatically search for the optimal scheduling scheme with respect to different DCNN models for energy efficiency. Finally, we conduct an empirical study showing that 9Ring can reduce energy consumption by 31.4% and 43.9% on average, and improve performance by 12% and 10% on average, compared with Tetris and the NN accelerators on conventional low-power DRAM memory systems, respectively.