FPUGen: A FrameWork to Generate Custom Floating Point FMA Accelerators on FPGAs

Himanshu Rai, Aishwarya Sridhar, Wolfgang Ecker, Nanditha Rao · 2025

Floating-point representations enable finer granularity of computations, improving the overall accuracy of ML models. Approximate computing techniques using variable bit-width can be employed to reduce the computational complexity of training algorithms while maintaining an acceptable level of accuracy. In this work, we propose FPUGen: a framework to generate custom floating-point (FP) fused multiply-add (FMA) accelerators. The accelerators have several architectural highlights such as: (a) Reconfigurable FP-FMA units for any mantissa and exponent size, (b) Dynamic precision FMA (DFMA) support for one SP-32, two TF-32, two HP-16, or three BF-16 operations in parallel, (c) Rounding logic for dynamic precision and (d) pipelined/non-pipelined FP-FMA architecture support. We note that choosing BF-16 over DP-64 and SP-32 saves energy by 5.73x and 2.56x respectively on an FPGA. Our proposed DFMA-V2 architecture is 1.49x more resource efficient, has 3.85x better throughput, and is 1.71x more energy efficient compared to the baseline design (DFMA-V1) due to the pipelined and resource-sharing architecture. It achieves a maximum throughput of 56.10 GFLOP/s on the ZCU104 FPGA, which is 5.8x higher than the baseline design.

Read the paper · More papers on PaperTik