LU Decomposition on GPUs: The Impact of Memory Access
Leandro Fontoura Cupertino, Anderson Pires Singulani, Cleomar P. da Silva, Marco Aurélio C. Pacheco, Ricardo Farias · 2010
Graphics Processing Units (GPUs) are emerging as an attractive computing platform for general purpose computations due to their extremely high floating-point processing performance and their comparatively low cost. In the context of dense linear algebra, the LU decomposition represents a fundamental step in many computationally intensive scientific applications. The use of GPUs can accelerate the computation many times the speed of a single CPU. In this work, we investigate different implementations of the LU decomposition algorithm in a GPU. Our main goal is to parallelize the LU decomposition to fit the highly parallel architecture of modern GPUs, and to evaluate different types of memory access and their impact on the execution time of the algorithm. The results demonstrate that the memory access pattern can significantly impact the performance of the GPU implementation.