Optimising GS2 through Hybrid Parallelisation
David J. Hardman · 2013
Clustered architectures are now ubiquitous in high performance computing; nearly all supercomputers today are built using a collection of nodes, each having multiple cores with shared memory. The hybrid programming model has been born out of the need for a programming model that fits this hybrid architecture; suited to both the shared and distributed memory aspects of such a system. It is debated, however, whether or not hybrid implementations lead to better performance than simply having a messagepassing implementation. GS2 is a gyrokinetic simulation code parallelised using the MPI message-passing library. In this project we investigated the addition of OpenMP directives to GS2, in order to create a hybrid version of the code. We then set about testing the performance of our hybrid version compared to that of the original MPI-only version. It was found that with the correct ratio of MPI processes to OpenMP threads, a performance increase is possible. We found this performance increase to be cumulative and proportional to the number of time steps simulated, i.e. the more time steps simulated the larger the increase in performance when compared to the MPI-Only version of GS2. Furthermore, the addition of OpenMP threads allowed GS2 to scale to much larger core counts, as well as outperforming the like for like MPI process counts.