LU Factorization with Partial Pivoting for a Multi-CPU, Multi-GPU Shared Memory System
Lawrence Berkeley National Lab. (LBNL), Berkeley, CA (United States), Jakub Kurzak, USDOE Office of Science (SC), Pitior Luszczek, Mathieu Faverge, Jack J. Dongarra · 2012
LU factorization with partial pivoting is a canonical numerical procedure and the main component of the High Performance LINPACK benchmark. This article presents an implementation of the algorithm for a hybrid, shared memory, system with standard CPU cores and GPU accelerators. Performance in excess of one TeraFLOPS is achieved using four AMD Magny Cours CPUs and four NVIDIA Fermi GPUs.