Optimizing Off-Chip Memory Access for Deep Neural Network Accelerator
Yong Zheng, Haigang Yang, Yi Shu, Yiping Jia, Zhihong Huang · IEEE Transactions on Circuits & Systems II Express Briefs · 2022
Off-chip memory, such as DRAM, its access energy cost is orders of magnitude higher than other operations such as multiply and accumulate, thereby dominating the system energy consumption. Therefore, optimizing the access of the off-chip memory is crucial to further improve the energy efficiency of the deep neural network (DNN) accelerator. Towards this, this brief proposed an adaptive scheduling algorithm to minimize the DRAM access. Compared with the previous works, it can not only dynamically determine the data partition and the data type that will be reused, but also considered the constraints between adjacent layer, that is, if the output feature map of$i$th layer is divided into N parts, the output feature map of$i+1$th layer can only be divided into N parts or write back to the off-chip memory. Therefore, a minimize and realizable memory access solution can be obtained. Choosing 3 popular networks UNet, VGG-16 and MobileNet as benchmarks, the experiment results show that our scheduling algorithm can achieve a$34.46\% \sim 93.42\%$reduction in energy consumption of DRAM access and a$34.34\% \sim 93.37\%$reduction in DRAM access latency when compared to previous works.