Accelerating Matrix Multiplication on FPGAs
Rasha El-Atfy, Mohamed Amin Dessouky, Hassan El-Ghitani · 2007
In this paper we present a new architecture for fixed-point matrix multiplication using Xilinx Virtex4 device. The architecture effectively utilizes the hardware resources on the entire FPGA and makes use of DSP blocks inside the FPGA devices. The architecture also reduces the routing complexity. Our architecture can be implemented for non-square matrix multiplication. The proposed implementation shows improvement in area and latency compared to recent published work. An improvement by over 50% in FMAX and 20% in area using new FPGAs has been achieved.