Improved SSOR and incomplete Cholesky solution of linear equations on shared memory and distributed memory parallel computers
Wayne D. Joubert, Thomas C. Oppe · Numerical Linear Algebra with Applications · 1994
Abstract In this paper two new implementations of SSOR and incomplete factorization preconditioners are given, for shared memory and distributed memory parallel computers respectively. These new implementations give increased solution speeds for matrix problems such as those arising from discretized partial differential equations with natural ordering of the grid points, for which it is well‐known that the standard implementation of these preconditioners is difficult to parallelize effectively. For shared memory machines, a new technique is presented here which decreases the number of synchronization points in each preconditioning step and thus allows better parallel speedups. For distributed memory machines, an implementation based on block cyclic reduction is given which circumvents the problem of idle processors during the preconditioning phase. Descriptions of the implementations are given, and numerical comparisons are given for a model diffusion problem on the Cray Y‐MP and the CM‐2 Connection Machine.