A Scalable ARM+FPGA-Based CNN Accelerator with Limited Hardware Resources
Jinlin Ye, Wei Zhang · 2023
General processors cannot meet the requirements of low power consumption and high performance in mobile application scenarios, the hardware research platform of CNNs has begun to turn to high-performance computing platforms such as GPUs and FPGAs. The GPU contains a large number of stream processors to execute operations in parallel, which greatly reduces the operation time of the model. However, because of its large power consumption, mobile applications can hardly afford the GPU's power requirement. Compared with GPUs, FPGAs also have parallel architecture, but the power consumption is much lower, which is suitable for CNN mobile application scenarios. In this work, concerning that FPGAs with large resource can meet the deployment requirements of CNNs but they are too expensive. How to deploy CNNs with little-resource has become an urgent problem to be solved. Based on this, we propose a scalable ARM+FPGA-based CNN accelerator with limited hardware resources, data flow is controlled through ARM, network is computed through FPGA. LeNet is adopted as the acceleration target, experiments were carried out on MNIST datasets and under the little-resource FPGA xc7z020clg400-2. The result show that at the frequency of 100 MHz, the accuracy of the accelerator prediction of 10000 images is 98.57% and the processing time of each image is 16.09ms.