Memory optimization of CNN Heterogeneous Multi-Core Architecture

Muxuan Gao, Dake Liu, Zhidong Miao, Shaohan Liu · 2019

Convolutional neural networks(CNNs) computing complexity raises challenges to the power and cost of CNN accelerators. The power consumption and delay of the data access between on-chip and off-chip memories have been a bottleneck of the on-chip system for CNN acceleration. We propose a static scheduling framework based on our heterogeneous programmable multi-core ASIP (Application Specific Instruction-set Processor) for inference of deep learning. This scheduling framework minimizes power consumption and promotes performance under the hardware resource limitation, including datapath computing capability, access bandwidth limitation, and on-chip memory size. We also discussed how to achieve theoretical maximum performance for a set of fixed hardware constraints. Comparing to the scheduling on DNA(deep neural architecture) under the same AlexNet platform, our scheduling achieves 11.8% more average resource utilization. By reusing and broadcasting local data, the average DRAM access of our design is reduced down to 41% of the access of DNA. The estimated power consumption of this design is 35% lower than the power of DNA, mostly because of the power reduction from DDR data accesses. Meanwhile, through the ping-pang memory and the careful design of data partition, the time of data access between DRAM and SRAM is reduced 77% by access hidden.

Read the paper · More papers on PaperTik