Benchmarking Floating Point Performance of Massively Parallel Dataflow Overlays on AMD Versal Compute Primitives

Mohamed Bouaziz, Suhaib A. Fahmy · 2025

Many massively parallel applications require floating-point (FP) precision, necessitating specialized hardware support. Novel Reconfigurable Dataflow Accelerators (RDAs) and Coarse Grained Reconfigurable Arrays (CGRAs) like the AMD Versal AI Engine Array can implement optimized datapaths. The AMD Versal FPGA family integrates AI Engine cores for vector FP operations as well as DSP58 primitives in the programmable logic for use in fine-grained architectures. Configurability at these levels involves factors, particularly data movement, that complicate measuring empirical compute performance limits. This paper presents an architectural model to isolate and measure these limits. Using this model, we compare resources, showing the superior operator density and performance of the DSP58 over the previous generation AMD UltraScale+ DSP48E2 for massively parallel dataflow overlays. We also show that the DSP58 can outperform programmable AI Engines in massively parallel feedforward applications.

Read the paper · More papers on PaperTik