On-chip Memory Optimized CNN Accelerator with Efficient Partial-sum Accumulation

Hongjie Xu, Jun Shiomi, Hidetoshi Onodera · 2020

In convolutional neural networks (CNNs), data movement inside convolution layers between memory and PEs is most energy dominant. This paper proposes a convolution processing dataflow that reduces both the number of memory accesses and the on-chip buffer capacity for convolution operations. Based on the dataflow, we design an on-chip buffer-minimized CNN accelerator. Compared with the state-of-the-art CNN accelerator, the proposed CNN accelerator utilizes 2.30 times less on-chip buffer and 2.18 times energy efficiency to achieve the same data throughput under Alexnet. The proposed architecture is able to achieve higher data throughput with the almost constant on-chip buffer capacity.

Read the paper · More papers on PaperTik