Optimization of Compiler-Generated OpenCL CNN Kernels and Runtime for FPGAs
Seung–Hun Chung, Tarek S. Abdelrahman · 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) · 2022
We translate frozen CNN models into OpenCL kernels with the TVM compiler and then use Intel's OpenCL SDK to compile to an FPGA bitstream. We improve the performance of the generated base hardware with optimizations that increase parallelism, reduce memory access latency, and save on-chip resources. We automate these optimizations in TVM and evaluate them by generating accelerators for LeNet-5, MobileNetV1 and ResNet-34 on an Intel Stratix 10SX. The optimizations improve the performance of the generated accelerators by up to 846 × over the base ones. The optimized accelerators are up to 4.57 × faster than TensorFlow on CPU, 3.83 × faster than single-threaded TVM and are only 0.34 × slower than TVM with 56 threads. Our optimized kernels also outperform ones generated by similar approaches that use high-level synthesis, but they underperform ones that utilize hand-optimized designs. Thus, our approach is most useful in environments that benefit from increased performance and fast prototyping, realizing the benefits of FPGAs without hardware design expertise.