Algorithmic and software development advances for next‐generation heterogeneous platforms
Roman Wyrzykowski, Florina M. Ciorba · Concurrency and Computation Practice and Experience · 2022
Heterogeneity is emerging as one of the most profound and challenging characteristics of today's and tomorrow's parallel and distributed computing environments, presenting new and exciting opportunities for their development. Most modern computing systems are heterogeneous, either for organic reasons because components grew independently, as is the case of desktop grids, by design to leverage the strength of specific hardware, as is the case of accelerated systems, or both. The impact of heterogeneity on all forms of parallel and distributed computing is increasing rapidly. Traditional algorithms, programming environments, and tools designed for legacy homogeneous systems will at best achieve a small fraction of the efficiency and the potential performance expected from parallel computing in tomorrow's highly diversified and mixed architectures. Innovative ideas, fresh models, novel algorithms, and other specialized or unified programming environments and tools are needed to efficiently use these new and increasingly diverse computing systems—for accelerating scientific discovery and impactful innovation. The International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar) has been the premier forum over the last 20 years, bringing together researchers to discuss these challenges and the solutions. The wide range of topics includes achieving performance portability on heterogeneous architectures, advances in software environments that facilitate efficient use of heterogeneous systems, performance and energy optimization of numerical and machine learning algorithms on heterogeneous platforms, to name a few. The works presented at the HeteroPar'2020 workshop covered topics clearly exhibiting the significance and growth of the heterogeneous computing field. However, one general trend is apparent: the broad adoption of Graphics Processing Units (GPU) accelerators. Over the last decade, GPUs have been established as the main powerhouse in leadership supercomputers and an invaluable component to accelerate computations for a vast spectrum of applications—from numerical linear algebra libraries powering computational science to various machine learning workloads. This trend is evidenced by the increasing number of GPU-related publications submitted to HeterPar and supported by growing diversity within the GPU world, where AMD accelerator architectures start to compete with Nvidia's comprehensive solutions, along with the third GPU accelerator option—from Intel—available soon. This special issue of Concurrency and Computation: Practice and Experience contains six selected papers from the HeteroPar'2020 workshop. We hope you find them interesting and stimulating new ideas and future advancements for next-generation heterogeneous platforms. The 18th International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar'2020) was held in Warsaw, Poland, on 25 August 2020. For the 11th time, this workshop was organized in conjunction with the Euro-Par annual series of international conferences. Because of the COVID-19 pandemic, HeteroPar'2020 was held as a virtual event. Sixteen articles were submitted for review, with authors from eight countries. Each paper secured at least three reviews from members of the program committee, whereas 12 submissions received at least four reviews. After a thorough peer-reviewing process that included discussion and agreement among reviewers whenever necessary, nine articles were selected for presentation at the workshop. The review process focused on the quality of the papers, their innovative ideas and applicability to heterogeneous computing. The topics addressed in the accepted papers include domain-specific languages for numerical algorithms, virtualization for CUDA applications, unified memory in CUDA, porting CUDA codes to AMD GPUs, management of heterogeneous cloud resources, GPU implementation of graph neural networks, GPU and CPU signal processing for a wildlife tracking system, parallelization of the k-means algorithm on CPU-GPU platforms, and a portable solver for systems of linear equations. An integral part of the workshop was two keynote talks given by Enrique S. Quintana-Orti (Technical University of Valencia, Spain) and Tal Ben-Nun (ETHZ Zurich, Switzerland) devoted to using approximate and transprecision computing in sparse linear solvers, and data-centric approach for performance portability on heterogeneous architecture, respectively. After the workshop, the program committee invited the authors of the presented works to submit revised and extended versions of their contributions as part of the papers submitted to this special issue. These new versions were reviewed independently again by at least three reviewers. Finally, six papers were accepted for publication in the special issue. They are summarized below. Aliaga et al.1 focus on optimizing the sparse matrix–vector product (SpMV), which dictates, to a large extent, the performance of a considerable variety of scientific applications. The proposed approach introduces a variant of the coordinate sparse matrix format that allows combining load-balancing with compressing both the indexing arrays and the numerical information to reduce the pressure on memory while using the available compute power of modern CPUs and GPUs efficiently. This approach is multi-platform, in the sense that the realizations are built upon common principles but differ in the implementation details, which are adapted either to avoid thread divergence in the GPU case or to maximize compression for multicore architectures. The evaluation on the two last generations of NVIDIA GPUs as well as Intel and AMD processors demonstrate the benefits of the new kernels compared with the optimized implementations of SpMV in Nvidia's cuSPARSE and Intel's MKL libraries. k-Means is a standard algorithm for clustering data used as the final step for high-quality spectral clustering. To overcome the scalability challenge when processing large datasets, the authors of paper2 propose to apply also the k-means algorithm as a preprocessing task to reduce the input data instances. Additionally, parallel optimization techniques are introduced to improve the efficiency of the k-means algorithm on CPU and GPU. Notably, a two-step summation method with package processing is used to handle the effect of rounding errors that may occur during the phase of updating cluster centroids. The extensive experiments on synthetic and real-world datasets containing millions of instances exhibit a speedup up to 7 for the k-means iteration time on GPU versus 20/40 CPU threads using AVX units while achieving double-precision accuracy with single-precision computations. Dmitruk et al.3 show how the OpenACC standard can be efficiently used to implement solvers for tridiagonal Toeplitz systems of linear equations for a variety of modern GPU-accelerated and multicore architectures. Two parallel algorithms are studied concerning particular assumptions about coefficient matrices. In the first case, a new, faster implementation of the divide and conquer method is proposed, while in the second one, a novel, vectorizable algorithm is introduced. Using both column-wise and row-wise matrix storage formats is studied, along with efficient conversion between them using cache memory to improve the overall performance. It is also shown how to tune the performance by predicting the best values of the methods' parameters. Numerical experiments performed on Intel CPUs and Nvidia GPUs confirm the excellent performance and accuracy of the developed implementations. Robust high-performance implementations of signal-processing tasks performed by a high-throughput wildlife tracking system are presented by Rubinpur et al.4 The system tracks radio transmitters attached to wild animals by estimating the time of arrival of radio packets to multiple receivers. The time-consuming estimation of wideband radio signals is a bottleneck that limits the system's throughput. A sequential high-performance CPU implementation has been developed first, and then a GPU implementation to overcome this bottleneck. The authors carefully evaluate the performance of these real-world codes. The evaluation indicates that the GPU version dramatically improves both performance and power-performance efficiency relative to a desktop CPU—a scenario typical for current base stations. Performance improves by more than 50 times on a high-end GPU and more than four times with a GPU platform that consumes almost five times less power than the CPU one. The desire to take advantage of virtualization in heterogeneous computing resources with GPU accelerators motivates Eiling et al.5 Currently, GPUs do not offer virtualization support that enables fine-grained control, increased flexibility, and fault tolerance. The authors present Cricket—a transparent and low-overhead solution to GPU virtualization that enables future research of various virtualization techniques, due to its open-source nature. Cricket supports remote execution and checkpoint/restart of CUDA applications. Both features allow the distribution of GPU tasks dynamically and flexibly across computing nodes and the multitenant usage of GPU resources, improving their flexibility and utilization in high-performance and cloud computing. Solving partial differential equations (PDEs) on unstructured grids is a cornerstone of engineering and scientific computing. Alhaddad et al.6 introduce the HighPerMeshes C++-embedded domain-specific language (DSL) that bridges the abstraction gap between the mathematical formulation of mesh-based algorithms for PDE problems and an increasing number of heterogeneous platforms with their various programming models. The HighPerMeshes DSL aims at higher productivity of the code development for multiple target platforms. For this aim, the OpenCL is used as a backend, targeting various GPUs and other heterogeneous architectures such as FPGAs. Apart from describing the basic structure of the DSL, its usage is demonstrated with three examples. The mapping of the abstract algorithmic description onto parallel hardware, including compute clusters, is also presented. Finally, the achievable performance and scalability are demonstrated for different example problems. The guest editors of this special issue wish to thank the authors of the submitted papers, the reviewers for the careful evaluation of the papers, and the valuable suggestions that helped the authors to improve their contributions. In addition, we would like to sincerely thank Prof. David W. Walker (Editor-in-Chief of Concurrency and Computation: Practice and Experience) for the opportunity to guest edit this special issue and for his guidance during this process. Data sharing is not applicable to this article as no datasets were generated or analyzed in this study.