Optimizing an Atomics-Based Reduction Kernel on OpenCL FPGA Platform

Zheming Jin, Hal Finkel · 2018

Field-programmable gate arrays (FPGAs) are becoming one of heterogeneous computing components in highperformance computing. To facilitate the use of FPGAs for developers and researchers, high-level synthesis tools are pushing the FPGA-based design abstraction from the register-transfer level to high-level language design flow using OpenCL/C/C++. Currently, there are few studies on the atomic functions in the OpenCL-based design flow on an FPGA. In this paper, we evaluate the performance of atomic functions using a reduction kernel on an OpenCL FPGA platform as a case study. We describe the implementations of an integer sum-reduction kernel in OpenCL, and perform the optimizations of memory accesses. Fully utilizing the bandwidth of the data bus can bring a factor of 15 improvement over the baseline kernel. The performance speedup of the kernel using local memory for atomic operations is 6.8X over the naïve kernel using global memory. The combination of both optimizations can lead to 112X speedup. Compute unit duplication can be applied to the kernel to further improve the performance by a factor of 2.9.

Read the paper · More papers on PaperTik