On Simplifying and Optimizing Programs for Heterogeneous Computing Systems
Ivan Grasso · Digital Library of the University of Innsbruck (University of Innsbruck) · 2017
Starting with OpenMP 4.0, it is possible to offload parallel regions to different devices, such as GPUs or accelerators.The OpenMP directives are platform independent, allowing a high degree of portability and performance with little programming effort.OpenACC [88] provides a collection of compiler directives to specify loops and regions of code in standard C, C++, and Fortran to be offloaded from a host CPU to an attached GPU or accelerator.The OpenACC directives allow programmers to create high-level host/device programs without the need to explicitly initialize the devices or manage data transfers.All of these details are implicit in the programming model and are managed by the OpenACC compilers and runtimes.Recently, SYCL [62], a new abstraction layer that builds on the concepts of portability and efficiency, was introduced.SYCL is a royaltyfree, cross-platform layer that enables code for heterogeneous processors to be written in a single-source style using standard C++.Template functions can contain both host and device code to construct complex algorithms that use OpenCL acceleration. thesis goals and organizationThis thesis will focus on OpenCL that currently offers a clear advantage in terms of number of accessible devices and level of hardware control.Using the OpenCL programming model, we will explore and discuss novel approaches aiming at both simplifying and optimizing programs for heterogeneous computing systems.This thesis is divided into five additional chapters as follows:Chapter 2 introduces the models that describe all the abstractions used in this work.Chapter 3 focuses on embedded systems analyzing the performance and energy advantages of embedded GPUs for HPC.We identify, implement and evaluate software optimization techniques for efficient utilization of the ARM Mali GPU Compute Architecture showing for the first time that embedded GPUs have qualities that make them good candidates for HPC systems.Chapter 4 investigates the distribution of tasks among the available devices in order to maximize the performance of heterogeneous com-