Solving “Large” Dense Matrix Problems on Multi-Core Processors and GPUs

Mercedes Marqués-Andrés, Gregorio Quintana‐Ortí, Enrique S. Quintana–Ort́ı, Robert A. Geijn · Repositori UJI (Universitat Jaume I) · 2009

Few realize that, for large matrices, many dense matrix computations achieve nearly the same performance \t\t\t\t when the matrices are stored on disk as when they are stored in a very large main memory. Similarly, few realize that, given \t\t\t\t the right programming abstractions, coding Out-of-Core (OOC) implementations of dense linear algebra operations (where \t\t\t\t data resides on disk and has to be explicitly moved in and out of main memory) is no more difficult than programming \t\t\t\t high-performance implementations for the case where the matrix is in memory. Finally, few realize that on a contemporary \t\t\t\t eight core architecture or a platform equiped with a graphics processor (GPU) one can solve a 100, 000 × 100, 000 \t\t\t\t symmetric positive definite linear system in about one hour. Thus, for problems that used to be considered large, it is not \t\t\t\t necessary to utilize distributed-memory architectures with massive memories if one is willing to wait longer for the solution \t\t\t\t to be computed on a fast multithreaded architecture like a multi-core computer or a GPU. This paper provides evidence in \t\t\t\t support of these claims

Read the paper · More papers on PaperTik