Parallel computing on graphics processing units and heterogeneous platforms
Paolo Bientinesi, José R. Herrero, Enrique S. Quintana–Ort́ı, Robert Strzodka · Concurrency and Computation Practice and Experience · 2014
This special issue contributes to the field of parallel computing on graphics processing units and heterogeneous platforms with extended versions of selected papers from two workshops, namely the 3rd Minisymposium on GPU Computing—held as part of the 10th International Conference on Parallel Processing and Applied Mathematics (PPAM 2013) in Warsaw, Poland—and the 11th International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar'2013)—held in conjunction with the Euro-Par 2013 conference in Aachen, Germany. During the past decade, high-performance computing evolved toward multi-core and many-core architectures. General-purpose processors feature now dozens of coarse-grain (complex) cores each with four to eight SIMD lanes for parallel computation and multi-channel memory buses for high bandwidth. Hardware accelerators such as graphics processing units (GPUs) have also a two-stage design with multiple coarse units that contain an even higher number of SIMD lanes and wider memory buses for higher bandwidth. Therefore, the adoption of hardware accelerators is rapidly advancing in performance sensitive areas. They are particularly relevant in high-throughput disciplines such as high-quality 3D computer graphics and vision, real-time data stream processing, and high-performance scientific computing. The main reason behind this trend is that these accelerators can potentially yield speedups and energy savings orders of magnitude higher than those obtained with optimized implementations for general-purpose CPU cores. A clear indicator of this trend is the prevalence of these accelerators in the supercomputing systems in the top positions of both the TOP500 and Green500 lists. As a result, during the past few years, these architectures have become powerful, capable, and inexpensive mainstream coprocessors, useful for a wide variety of applications. Furthermore, they are nowadays present in a large variety of machines, ranging from low-end single user-platforms to supercomputers. However, the benefits of heterogeneous systems do not come ‘for free’: scientists using these platforms have to deal not only with multiple parallelism levels, but also with the programmability differences of available accelerators. To address these challenges, we observe the development of a very rich environment for their programming, particularly in comparison with the restricted landscape of only a few years ago. A key criterion to characterize the new high-level programming tools and libraries for these devices is their positioning within the triangle of performance, coding comfort and specialization. The spectrum ranges from high-performance building blocks for common numeric or discrete transformations, to domain-specific libraries that facilitate the solution of a certain class of problems, and to general high-level abstractions targeted toward increasing programmers' productivity. In summary, the advances both in the hardware and in the programmability of accelerators, coupled with their potentially appealing performance/power ratio for a wide range of applications, have pushed organizations to invest in heterogeneous systems that include accelerators and have motivated researchers to port their algorithms to such systems and develop novel tools to facilitate their usage. This special issue contributes to this important field with extended and carefully reviewed versions of selected papers from two workshops, namely the 3rd Minisymposium on GPU Computing, which was held as part of the 10th International Conference on Parallel Processing and Applied Mathematics (PPAM 2013) in Warsaw, and the 11th International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar'2013), which was held in conjunction with the Euro-Par 2013 conference in Aachen. Nineteen papers were published in the Conference Proceedings of these two events, after one or two review rounds. Extended versions of selected papers went through two new review rounds, resulting in the acceptance of the nine papers contained in this special issue. The topics offer a good cross section of current challenges on heterogeneous computing: further abstractions in the programming model, advances in the scheduling of tasks or their communications, improvements of basic parallel algorithms in discrete mathematics and linear algebra, and the utilization of the parallel processing power of GPUs for real-world applications. In 1, the authors propose the design of a directory, along with a reduced runtime application binary interface, to handle data management between a host and accelerators in the OpenMP 4.0 and OpenACC standards. Some extensions were added to the directory to allow more flexibility when handling subarrays in the data clauses, including support for unstructured data lifetime. With these modifications, one can use multiple parts of the same array in a nested data environment, keeping the coherence in the accelerator memory between all the subparts. In 2, the paper addresses the solution of large-scale eigenvalue problems that appear in the motion simulation of complex macromolecules on multi-threaded platforms. They compare implementations of three high-performance eigensolvers using out-of-core techniques, enhancing their performance by leveraging hybrid CPU-GPU routines. They show that a Krylov subspace-based eigensolver presents a much lower theoretical cost and outperforms the GPU alternatives for macromolecular simulations. In 3, the authors describe how to conduct high-performance tracking of 3D human motion in real-time using multi-view images and particle swarm optimization. The tracking involves configuring the 3D human model, in the pose described by each particle, and then rasterizing it in each particle's 2D plane. Image acquisition and image processing are multi-threaded and run on CPU in parallel with particle swarm optimization-based searching which is GPU-accelerated, obtaining more precise tracking. In 4, an automated approach to estimate the memory footprint of non-linear data objects is presented. This is a novel method to build a graph-based static data type descriptions that allow to create code for injectable functions that automatically determine the memory footprint of data objects at run-time. This is useful in the context of current programming models for heterogeneous devices with disjoint physical memory spaces which require explicit allocation of device memory and explicit data transfers. This task becomes difficult for non-linear objects, for example, linked lists or multiple inherited classes, due to memory requirements known only at run-time and the composition of complex data structures from basic types. In 5, the authors derive an asymmetric network property on TCP layer for concurrent bidirectional communications on Ethernet clusters and develop a communication model to characterize the communication times accordingly. They show that if the asymmetric network property is excluded from the model, the communication time predictions will be significantly less accurate than those made by using the asymmetric network property. In 6, the authors examine the possibilities of using a GPU for complex 3D finite difference computation. Parallel simulation algorithms using shared and surface memory for relativistic hydrodynamics problems are implemented. Their main objective is to design an efficient algorithm that can benefit from the properties of surface memory optimized for 2D spatial locality and compare it to the best known approach working on shared memory. Their results expose that surface memory is a promising approach for complex 3D finite difference methods. In 7, we are concerned with the challenges underlying the translation of advanced magnetic resonance imaging protocols into a clinical environment. Specifically, rapid online reconstructions require significant computational power. The authors address this problem by developing an external, online, heterogeneous image reconstruction system for magnetic resonance data. The system integrates an external computer equipped with a GPU card into the magnetic resonance scanners image reconstruction pipeline. The system promotes fast online reconstruction for computationally intensive algorithms turning them feasible in a busy clinical service. In 8, the authors compare the performance of various algorithms for the reduction of collective operations in a non-clairvoyant setting, that is, when the algorithms are oblivious to the communication and computation costs. Communication times can rarely be predicted with high accuracy, and may vary significantly over time. The paper assesses how classical static algorithms, where the tree is built before the actual reduction, perform in such settings and quantifies the potential advantage of dynamic algorithms, where the tree is built at run-time and depends on the actual duration of the operations. The study includes both commutative and non-commutative reductions. In 9, a new method for scheduling efficiently parallel applications on hybrid architectures (multi-core machine with GPUs) with m CPUs and k GPUs is presented. Thereby each task of the application can be processed either on a core (CPU) or on a GPU. The objective is to minimize the maximum completion time (makespan). The corresponding scheduling problem is NP-hard, and the authors propose an efficient approximation of a generic methodology. The main idea of the approach is to determine an adequate partition of the set of tasks on the CPUs and the GPUs using a dual approximation scheme. We would like to thank the authors for their excellent contributions to this special issue. The anonymous reviewers who helped to greatly improve the quality of the papers deserve also special acknowledgement; without their selfless effort, this special issue would not have been possible. We hope that the here assembled body of work inspires future research in the area of parallel computing on GPUs and heterogeneous platforms.