Acceleration of LU decomposition supporting double-double, triple-double, and quadruple-double precision floating-point arithmetic with AVX2

Tomonori Kouya · 2021

In this paper, we report the results obtained from the acceleration of multi-binary64-type multiple precision LU decomposition with Intel's Advanced Vector Extensions 2 (AVX2). We targeted double-double (DD), triple-double (TD), and quad-double (QD) precision arithmetic designed using certain types of error-free transformation (EFT) arithmetic. We implemented accelerated DD, TD, and QD precision addition and multiplication using SIMDized EFT functions with AVX2, which perform simultaneous computation using four binary64 numbers on the x86_64 computing environment, and with these, we were able to develop multiple-precision LU decomposition based on SIMDized matrix multiplication. Our LU decomposition is up to three times faster than the non-accelerated one.

Read the paper · More papers on PaperTik