Low/Adaptive Precision Computation in Preconditioned Iterative Solvers for Ill-Conditioned Problems
Masatoshi Kawai, Kengo Nakajima · 2022
Double precision(FP64) is generally used in computer science, and an incomplete Cholesky preconditioned conjugate gradient(ICCG) method, which is widely used in computer simulations, is a typical example. Recently, the use of single-precision(FP32) in the ICCG method with well-conditioned problems is discussed for reducing computational time and power consumption. Also, the use of half-precision(FP16) is examined in the field of machine learning, and hardware support of FP16 on GPUs and some CPUs is advancing. Using FP16 can further reduce the amount of memory translation from FP32, so high effects can be expected if it can be used with the ICCG method. When we apply FP16 to the ICCG method, It is difficult to solve the ill-condition problem because of the poor expressiveness of the exponent of FP16. Therefore, in this study, we evaluate the usefulness of FP16 in the ICCG method with several conditions of the problem, and we also proposed the implementation for CPU of and evaluated FP21 and FP42 adaptive precision whose expressiveness is higher than FP16. Additionally, to reduce the calculation time by using low precision, it is necessary to maintain high memory width consumption by an effective SIMDization. Therefore, in this study, we also evaluate the performance of ELLPACK(ELL) and Sell-C-σ storage formats, which can efficiently use SIMDization.