Accelerating LU-Decomposition of Arbitrarily Sized Matrices on FPGAs
Maram Krishna Kumar, Ziaul Choudry, Suresh Purini · 2023
In this paper, we design and develop a hardware accelerator for computing the LU decomposition of an input matrix. Our accelerator consists of two simple linear arrays of Processing Engines (PEs), one on each of the two SLR regions of the FPGA. All the computations arising from the block LU decomposition are simplified and scheduled on these two PE arrays. On an Alveo U50 FPGA, our design achieves a peak floating-point performance of 128 GLOPS/s and an average performance of 95 GFLOPS/s. We achieve ≈ 15× speedup on latency compared to an Intel MKL implementation on a 4-core Intel Xeon CPU.