Design Study on Impact of Memory Access Parallelism for Cloud FPGAs
Arnab A Purkayastha, Hamed Tabkhi · 2021
OpenCL parallelism capabilities combined with FPGAs deep pipelined data-path have largely improved the high-performance execution of massively parallel applications in the High-Performance Computing (HPC) environment. Despite this, memory access latency is an important bottleneck that affects performance. This paper presents an in-depth understanding of two optimization techniques, Double Data Rate (DDR) and Burst Transfer [BT], and formalizes generic framework(s) to improve memory access parallelism. We apply DDR and BT across a variety of applications on the AWS cloud-based Xilinx VU9FP FPGA with minimum programming complexity. Our results show that FPGA-aware OpenCL codes with DDR optimization result in an average speedup of 1.4X times over baseline with minimal power and utilization overheads. Similarly, BT results in a 1.5X improvement in performance and a 5% rise in bandwidth utilization. For performance comparison, we also stack up best-case FPGA numbers against naive CPU and GPU implementations on all the applications and observe that FPGA beats CPU and GPU numbers for at least 4 of the 9 applications.