Enabling high-level parallel programming on multi-FPGA clusters

Juan Miguel de Haro Ruiz, Carlos Álvarez, Daniel Jiménez-González, Xavier Martorell · 2024

Field Programmable Gate Arrays (FPGA) are still relatively new in the High Performance Computing (HPC) field. Hence, they still lack a mature ecosystem that allows non-FPGA experts to scale an application with many devices operating in parallel. In this paper, we add support for message passing inspired by the Message Passing Interface (MPI) to the Marenostrum Exascale Emulation Platform (MEEP) cluster, a state-of-the-art FPGA cluster. We use the OmpSs@FPGA programming model, which allows C/C++ code to run on the FPGA and call functions that behave like the well-known MPI_Send/Recv. For that, we implement the message passing runtime over the MEEP 100Gb Ethernet network. This network includes a switch connected to the QSFP port of the FPGA cards. The switch enables all-to-all connectivity without adding routers in the FPGA fabric. We also introduce a method to manage FPGAs that are PCIe-hosted by remote CPU nodes. I.e. from any CPU node we can load the bitstream, configure it, and transfer data to FPGAs that are attached to a different node. Finally, we evaluate the bandwidth of FPGA-FPGA, and CPU-FPGA local and remote communication, as well as the performance of benchmarks scaling from 1 to 64 FPGAs using the infrastructure presented in this paper. The benchmarks are N-body, Heat with Gauss-Seidel solver and Cholesky. We compare the results with the MareNostrum 4 supercomputer and get 2.3x and 3.5x better performance per power for the N-body and Heat.

Read the paper · More papers on PaperTik