Acceleration of in-core LU-decomposition of dense MoM matrix by parallel usage of multiple GPUs
Branko Lj. Mrdakovic, Milan M. Kostić, Dragan I. Olćan, Branko M. Kolundžija · 2017 IEEE International Conference on Microwaves, Antennas, Communications and Electronic Systems (COMCAS) · 2017
Acceleration of in-core LU decomposition by using multiple GPUs in parallel is presented in this paper. Memory limitations of GPUs are overcome by using block LU decomposition, where the entire system matrix is stored in CPU RAM, while only currently processed blocks are stored in GPU VRAM. The presented algorithm for LU decomposition enables utilization of an arbitrary number of GPUs. High efficiency of the parallelization is achieved by specifically tailored load balancing over the GPUs in the most time consuming parts of the block LU decomposition. Presented results show that a dense MoM matrix with 100 000 complex unknowns in single precision can be LU decomposed in about 8 minutes, on a personal computer equipped with 8 low-cost GTX 680 GPUs.