VFMA: Scalable Floating-Point Accelerator for Vector FMA on FPGAs

Himanshu Kumar Rai, Sasi Snigdha Yadavalli, Aishwarya Sridhar, Nanditha Rao · 2025

Floating-point units (FPUs) provide a flexible approach to numerical computations, particularly in applications requiring varying levels of precision. Operations such as Fused Multiply-Add (FMA) have been integrated into processor instruction sets to optimize complex arithmetic tasks. In this paper, we introduce a Vector Floating-Point FMA (VFMA) accelerator on an FPGA with reconfigurable lane width and precision, allowing custom exponent and mantissa configurations. Additionally, we propose a Tiled VFMA (TVFMA) architecture to maximize the throughput on the AMD-Xilinx ZCU104 FPGA. A key observation is that DSP-based VFMA designs deliver higher throughput at higher precisions (SP-32, DP-64), while LUT-based designs offer comparable performance at lower bit precisions (BF-16, QP-8). The VFMA accelerator with an optimal lane width of vector-6 achieves a 5.56 x higher throughput than its non-vector counterpart. The proposed TVFMA accelerator with vector-6 achieves an impressive throughput of $633.2 \mathrm{GFLOP} / \mathrm{s}$ for QP-8 precision. Our work demonstrates on average 2.51x, 5.72x, and 4.37x higher throughput for HP-16, SP-32, and DP64, respectively, compared to related works.

Read the paper · More papers on PaperTik