HISQ inverter on Intel Xeon Phi and NVIDIA GPUs
Swagato Mukherjee, Olaf Kaczmarek, Christian Schmid, Patrick Steinbrecher, Mathias Wagner · Proceedings of The 32nd International Symposium on Lattice Field Theory — PoS(LATTICE2014) · 2015
The runtime of a Lattice QCD simulation is dominated by a small kernel, whichcalculates the product of a vector by a sparse matrix known as the "Dslash"operator. Therefore, this kernel is frequently optimized for various HPCarchitectures. In this contribution we compare the performance of the IntelXeon Phi to current Kepler-based NVIDIA Tesla GPUs running a conjugate gradientsolver. By exposing more parallelism to the accelerator through invertingmultiple vectors at the same time we obtain a performance 250 GFlop/s on botharchitectures. This more than doubles the performance of the inversions. Wegive a short overview of both architectures, discuss some details of theimplementation and the effort required to obtain the achieved performance.