A gentle introduction to GPU programming: conference tutorial
Eduardo Colmenares, Amy Knowles · Journal of computing sciences in colleges · 2017
Initially HPC was limited to a single core per node, eventually evolving to multicore CPUs, during all that time the computational power of multicore architectures was harnessed either by making use of distributed memory programming (MPI) and/or shared memory (OpenMP, Pthreads). In recent years, and thanks to the development of many-core GPUs, things have radically improved in the HPC arena, thanks to its outstanding computational capabilities. In todays world, a high percentage of the solutions to the most computational challenging problems owe their fast, innovate, and accurate solution to the correct balancing between multicore CPUs and many-core GPUs (Nvidia Products). GPU-accelerated computing leverages the parallel processing capabilities of GPU accelerators and enabling software to deliver dramatic increases in performance for scientific, artificial intelligence, deep learning, graphics, engineering, and other demanding applications. This tutorial will cover architectural concepts associated with Nvidia GPUs, as well as programming concepts. During the first part of this tutorial, multiple simple and easy to follow examples will be presented to highlight the mapping from hardware to programming, building important concepts that will allow the audience to understand to a more complex example (Matrix Multiplication).