Design of a Coarse-Grained Processing Element for Matrix Multiplication on FPGA

Yuichi Okuyama, Shigeyuki Takano, T. Shirai · 2014

In this paper, we discuss and evaluate about a grain size of the PE of a matrix operation specific architecture with fused multiply add (FMA) units, Rapid MatriX, on FPGAs. Recent FPGAs have many DSP blocks which are high-performance arithmetic units. Hereby, implementing functional units for matrix operation to array structure of the Rapid MatriX, we propose to use DSP blocks efficiently by increasing grain size of FMA unit. We implement the Rapid MatriX using the refined PEs on an FPGA. In addition, we evaluate the clock frequencies and the clock cycles of calculation. As a result, throughput of the PE for 4times 4 matrix FMA is 3.14 times in comparison with the original PEs of scalar FMA for 8times 8 matrix multiplication.

Read the paper · More papers on PaperTik