Distributed memory matrix-vector multiplication and conjugate gradient algorithms

John Gregg Lewis, Robert A. Geijn · 1993

The critical bottlenecks in the implementation of the conjugate gradient algorithm on distributed memory computers are the communication requirements of the sparse matrix-vector multiply and of the vector recurrences. The data distribution and communication patterns of five general implementations whose realizations demonstrate that the cost of communication can be overcome to a much larger extent than is often assumed are described. The results also apply to more general settings for matrix-vector products, both sparse and dense.

Read the paper · More papers on PaperTik