A Fast Floating-Point Multiply–Accumulator Optimized for Sparse Linear Algebra on FPGAs
Kun Li, Xiangyu Hao, Zhenguo Ma, Feng Nan Yu, Bo Zhang, Qianjian Xing · IEEE Transactions on Very Large Scale Integration (VLSI) Systems · 2025
This brief presents a pipelined floating-point Multiply–Accumulator (FPMAC) architecture designed to accelerate sparse linear algebra operations. By designing a lookup-table-based 5–3 carry-save adder (CSA) and combining it with a 3–2 CSA, the proposed design minimizes the critical path and boosts operational speed. Moreover, the proposed architecture takes advantage of data characteristics in sparse linear algebra to displace the shift unit in the critical accumulation loop, further increasing the throughput rate. In addition, the integration of a lookup-table-based leading-zero anticipator (LZA) enhances normalization efficiency. Experimental results show that, compared with reported FPMAC designs, the proposed architecture may achieve a significantly higher maximum clock frequency for single-precision floating-point operations.