Scalable Hybrid Loop- and Task-Parallel Matrix Inversion for Multicore Processors

Sandra Catalán, Francisco D. Igual, Rafael Rodríguez‐Sánchez, Enrique S. Quintana–Ort́ı · 2021

We propose a hybrid parallelization scheme for matrix inversion on multicore processors that combines a look-ahead technique to extract task-parallelism, at a high level, with loop-level parallelism to ensure an efficient utilization of the processor memory subsystem. As a result, our scheme outperforms the conventional approach for dense linear algebra operations, which simply extracts parallelism from a multi-threaded instance of the BLAS (Basic Linear Algebra Subprograms), but also the alternative based on OpenMP task-parallelism only, which is supposed to offer higher scalability. We provide an extensive collection of experiments supporting our remarks on two recent Intel- and ARM-based architectures, with a very large count of cores.

Read the paper · More papers on PaperTik