Accelerating 128-bit Floating-Point Matrix Multiplication on FPGAs
Fumiya Kono, Naohito Nakasato, Maho Nakata · 2023
General Matrix Multiplication (GEMM),$C=\alpha AB+\beta C$with input matrices$A, B$and scalar parameters$\alpha,\beta$, is widely used in scientific and engineering computations. Its precision is critical in determining the accuracy of the target applications. The requirements for the number of bits to represent floating-point (FP) numbers depend on individual applications. As the IEEE 754 standard [1] defines, FP formats and arithmetic are available in various precision, such as binary32 and binary64. Operations with higher precision like binary128 are desired for specific applications represented by Semidefinite Programming (SDP), a natural extension of linear programming that aims to minimize linear functions subject to certain constraints. Since binary128 is hardly supported as hardware, the performance of applications relying on its arithmetic is typically 100 to 1000x slower than that only relying on binary64. Therefore, acceleration of binary128 arithmetic is crucial to solve SDP and applications demanding binary128 fast.