A driver-based approach for DMA transfer between FPGA-GPU

Shun Kasai, Yasunori Osana · 2022

Many recent HPC clusters employ GPUs as accelerators to improve their performance per watt, but the power wall is still there. The use of FPGAs draws attention to achieving more power efficiency because the users can design their application-specific custom accelerator. The hybrid approach with GPUs and FPGAs is also a hopeful solution. Still the data transfer between FPGA and GPU is a significant bottleneck because there is no native support for direct FPGA-GPU DMA transfers. This paper shows the implementation and evaluation of two direct FPGA-GPU data transfer methods. By mapping FPGA memory on the BAR address space, the GPU's DMA gained only 300 MB/s, which was slower than indirect transfer through the host's main memory. The second method maps GPU memory on the BAR address space by NVIDIA GPUDirect API, and the FPGA's DMA initiates data transfer. Although GPU-to-FPGA data transfer is not still succeeded with this method, the FPGA-to-GPU bandwidth on the Gen3 x4 bus reached 3.25 GB/s: more than 80% of peak bandwidth. This result outperformed FPGA-Host-GPU transfer by 1.6x, even faster than Host-GPU DMA transfer. Since the second method is implemented by a small extension on Xilinx's XDMA driver, it is expected to be portable for other FPGA DMA controllers.

Read the paper · More papers on PaperTik