Scalable scientific computing applications for GPU-accelerated heterogeneous systems
Riesinger, Christoph · mediaTUM – the media and publications repository of the Technical University Munich (Technical University Munich) · 2017
In the last decade, graphics processing units (GPUs) became a major factor to increase performance in the area of high performance computing.This is reflected by numerous examples of the fastest supercomputers in the world which are accelerated by GPUs.GPUs are one representative of many-core chips, the major technology development to boost hardware performance in previous years.Many-core chips were one factor besides others such as progress in modeling, algorithmics, and data structures to allow scientific computing advance to its current state-of-the-art.To exploit the whole computational power of GPUs, numerous challenges have to be tackled in the area of parallel programming: Latest developments in handling the core characteristics of GPUs such as two additional levels of parallelism, the memory system, and offloading shift the focus on the usage of multiple GPUs in parallel and/or combining them with the performance of CPUs (heterogeneous computing).As a consequence, hybrid parallel programming (e.g.MPI, OpenMP) concepts are required and load balancing as well as communication hiding become even more relevant to achieve good scalability.In this work, we present approaches to benefit from GPUs for three different applications, each covering different algorithmic characteristics: First, a pipelined approach is used to determine the eigenvalues of a symmetric matrix not only enabling very high FLOPS rates but also allowing for the handling of even large systems on one single GPU.Second, the solution of random ordinary differential equations (RODEs) offers multiple levels of parallelism which is predestined for systems with multiple GPUs leading to the first implementation of an RODE solver to deal with problems of reasonable size.Finally, it is shown that a pioneering hybrid implementation of the lattice Boltzmann method making use of all available compute resources in the system where the CPU can process regions of arbitrary volume can attain good scalability.iii Even if there is only one author name written on the front page of this thesis, there are numerous other persons who contributed to this document in one way or the other.So I am taking the chance to express my acknowledgements and thanks to these people.First of all, I would like to mention my PhD supervisors Prof. Hans-Joachim Bungartz and Prof. Takayuki Aoki.They offered me the opportunity to start and successfully work on my PhD in very comfortable, pleasant, productive and competent environments, especially during my research stay abroad in Tokyo where the first actual results could be achieved.Before achieving actual results, much groundwork has to be finished, not always done by myself.Here, I want to thank Tobias Neckel and Florian Rupp for their preliminary studies in the field of random ordinary differential equations forming the theoretical basis of part III of this document.Special acknowledgements go to Tobias who did not just contribute in a technical way as the advisor of my thesis but also became a close friend.The same gratefulness belongs to Martin Schreiber and Arash Bakhtiari for their practical effort in the area of the lattice Boltzmann method continued by me in part IV.Martin, I am not sure if you reduced or actually extended the time to finish my PhD, anyways, you definitely made this time much more valuable.In addition, I would like to thank these people who gave me access to the computing resources essential for my research work.Robert Speck paved the way to utilize the infrastructure in Jülich, Frank Jenko enabled the access to the Max Planck resources in Garching, Maria Grazia Giuffreda arranged the usage of several clusters in Lugano, and, again, Prof. Takayuki Aoki is mentioned for his support in Tokyo.If there was any onside technical problem, Roland Wittmann was the guy you can count on.Furthermore, I have to thank