Exploiting Parallelism on Irregular Applications Using the GPU
Manuel Ujaldón, Joel Haskin Saltz · 2005
The computational speed on microprocessors is increasing faster than the communication speed, especially on parallel processors such as GPUs. Thus, the computations that benefit the most from GPU processing have high arithmetic intensity. This paper compares the effectiveness of GPUs when handling scientific general-purpose irregular problems, outperforming counterpart CPUs by a wide margin and identifying the AGP bus as the major bottleneck in graphics architecture. We study the impact that the emerging PCI-Express bus has for accelerating such applications when replacing AGP. A number of software optimizations are also conducted by using recent APIs, OpenGL extensions and drivers, leading to loading times 40 % lower on PCI-Express and four times faster when overlapping communication with computation. Execution times are shown on a benchmark composed of an Euler solver and a sparse matrix-vector product running on Nvidia GeForce FX and GeForce 6800 GT graphics cards. 1.