Economical Two-fold Working Precision Matrix Multiplication on Consumer-Level CUDA GPUs
Noriyuki Fujimoto · 2011
Dot product faithfully rounded after "as if" computed in K-fold working precision (K ≤ 2) is known to be computable only with floating-point numbers defined in IEEE 754 floating-point standard. This paper presents a CUDA GPU implementation of two-fold working precision matrix multiplication based on the dot product computation method. Experimental results on a GeForce GTX580 and a GTX560Ti show that the proposed implementation has 1.84 to 1.95 times higher GFLOPS performance in two- fold working precision compared to the performance of CUBLAS dgemm in double-precision on a Tesla C2070 high-end GPU. The proposed implementation can be used to obtain higher performance in pseudo double-precision with low cost consumer-level GPUs whose double-precision native performance is limited.