Special Issue: GPU computing

José R. Herrero, Enrique S. Quintana–Ort́ı, Robert Strzodka · Concurrency and Computation Practice and Experience · 2011

The combined hurdles of power consumption, limited instruction-level parallelism, and memory latency have led hardware manufacturers to design power aware multi-core processors and specialized many-core hardware accelerators in order to further exploit the increasing number of transistors dictated by Moore's Law. The race is open: As of today, Intel's top-of-the-line designs include 8 cores in its Xeon Nehalem architecture, AMD raises this number to 12 cores in the Opteron Magny-Cours processor, and the road-maps of the two companies indicate that these quantities will increase to 10–12 (Intel) and 16 (AMD) cores in 2011. At the same time, specialized hardware architectures, such as graphics processing units (GPUs), with tens of cores are already widely deployed. Core numbers are only one level of parallelism that is on the rise. Each core in current processors contains multiple processing elements enabling parallel processing within it. Power efficiency requires that the parallel processing elements are assembled in SIMD (Single Instruction Multiple Data) units. In the current CPU cores SSE instructions are supported enabling up to 4 parallel multiply-add operations on single precision floating point numbers. Soon this figure will rise to 8 with the AVX instruction set and the Intel Larrabee design featured already 16-wide SIMD units in each core. The number of instructions that can be executed in parallel on each GPU core varies strongly between 16 and 80 because of the different designs of GPU cores by different manufacturers; e.g. NVIDIA Fermi architecture has up to 16 cores times 32 processing elements, whereas AMD Cypress architecture contains up to 20 cores times 80 processing elements, in both cases 2 × increase over the previous generation. However, counting the processing elements on a GPU allows only a comparison within the same GPU family. Across families at least the differing factors of shader clock and arrangement of the processing elements must be taken into account. Although these new multi-core and many-core architectures can potentially deliver a revolutionary boost in raw performance, the efficient utilization of the growing SIMD and many-core parallelism is the key that will determine their success or failure. In this line, the recent advances in the hardware, functionality, and programmability of graphics processors (GPUs) have greatly increased their appeal as add-on co-processors for general-purpose computing. With the involvement of the largest processor manufacturers, NVIDIA, AMD, and Intel, and the strong interest from researchers of various disciplines, this approach has moved from a research niche to a forward-looking technique for heterogeneous parallel computing. Scientific and industry researchers are constantly finding new applications for GPUs in a wide variety of areas, including image and video processing, molecular dynamics, seismic simulation, computational biology and chemistry, fluid dynamics, weather forecast, computational finance, quantum physics, and many others. GPU hardware has evolved over many years from graphics pipelines with many heterogeneous fixed-function components over partially programmable architectures toward a more homogeneous general-purpose design (though some fixed-function hardware has remained because of its efficiency). The general-purpose computing on GPU (GPGPU) revolution started with programmable shaders. NVIDIA Compute Unified Device Architecture (CUDA) and, to a smaller extent, AMD CAL/Brook+ have brought GPUs into the mainstream of computing, developing what has been recently coinedGPU computing* . The great advantage of CUDA is that it defines an abstraction that presents the underlying hardware architecture as a sea of hundreds of fine-grained computational units with synchronization primitives on multiple levels. With OpenCL, there is now also a vendor-independent high-level parallel programming language and an application programming interface that offers the same type of hardware abstraction. GPUs are very versatile accelerators because besides the high hardware parallelism they also feature a high bandwidth connection to dedicated device memory. The latency problem of DRAM is tackled via a sophisticated thread scheduling and switching mechanism on-chip that continues the processing of the next thread as soon as the previous stalls on a data read. These characteristics make GPUs suitable for both compute- and data-intensive parallel processing. All together, the advances in the GPU hardware, the improvements in their programmability, as well as a potentially appealing performance/power ratio, have pushed organizations to invest in heterogeneous systems that include GPUs, and have motivated researchers to port their algorithms to such systems. This special issue contains extended versions of selected papers from the Minisymposium on GPU Computing, which was held as part of the Eight International Conference on Parallel Processing and Applied Mathematics—PPAM 2009 in Wroclaw (Poland). Ten papers were published in the Conference Proceedings, after two review rounds. Extended versions of some of these papers went through a new review process, resulting in the selection of papers contained in this special issue. The topics offer a good cross-section of the current GPU challenges: further abstraction of the hardware and the programming model, improvements of basic parallel algorithms in discrete mathematics and linear algebra, and the utilization of the parallel processing power of GPUs for real-world applications. Michael Repplinger and Philipp Slusallek (‘Stream processing on GPUs using distributed multimedia middleware’, Concurrency and Computation: Practice and Experience [this issue]) introduce an open distributed middleware for the development of applications in multi-GPU systems. In particular, the solution contributed by the authors can seamlessly integrate processing components, hide architecture-specific issues, combine GPUs and CPUs in a heterogeneous computational system, and use local and remote GPUs for distributed processing. Hagens Peters et al. (‘Fast in-place, comparison-based sorting with CUDA: a study with bitonic sort’, Concurrency and Computation: Practice and Experience [this issue]) present their work on a comparison-based in-place implementation of bitonic sort on CUDA-enabled GPUs. They identify and minimize the access to global memory as the main bottleneck and obtain remarkable sorting rates for a large number of sorting elements. Paolo Bientinesi et al. (‘Condensed forms for the symmetric eigenvalue problems on multithreaded architectures’, Concurrency and Computation: Practice and Experience [this issue]) analyze an alternative blocked algorithm for the solution of symmetric eigenvalue problems that can be efficiently cast in terms of efficient matrix–matrix products that attain high performance on a graphics processors. The experimental study of the authors using an accelerated version of this algorithm on NVIDIA GT200 generation of graphics processors demonstrates its superior performance compared with the traditional Level-2 BLAS-based approach on Intel Xeon E5520 (Nehalem) and E7640 (Dunnington) processors. Bernardo Rocha et al. (‘Accelerating cardiac excitation spread simulations using GPUs’, Concurrency and Computation: Practice and Experience [this issue]) employ a graphics processor to significantly accelerate the simulation of electrical activity in the heart. The authors' experiments with the solution of the ordinary differential equations modeling 2D cardiac tissues on a NVIDIA GeForce GT200 show a performance acceleration of 20–180 times with respect to a Quad-core processor. We thank the authors for their excellent contributions to this special issue as well as the anonymous reviewers who helped the authors and the editors of this issue to greatly improve the quality of the papers. We hope that it inspires future research in the area of GPU computing.

Read the paper · More papers on PaperTik