dCUDA: GPU Cluster Programming using IB Verbs
Lukas Kuster · Repository for Publications and Research Data (ETH Zurich) · 2017
Over the last decade, the usage of GPU hardware to accelerate high performance computing tasks increased due to the parallel computing nature of GPUs and their floating point performance.Despite the popularity of GPU accelerated compute clusters for high performance applications like atmospheric simulations, no unified programming model for GPU cluster programming is established in the community.dCUDA, a unified programming model for GPU cluster programming designed at ETH Zürich, is a candidate to fill this gap.dCUDA provides a device-side library for Remote Memory Access (RMA) of a global address space and supports fine-grained communication and synchronisation on the level of CUDA thread blocks.The dCUDA programming model enables native overlap of communication and computation for latency hiding and better utilisation of GPU resources by oversubscribing the system.dCUDA provides a well designed communication library and outperforms traditional GPU cluster programming approaches.We extend the dCUDA library and improve the performance of the dCUDA programming model.New functionality enables new low latency synchronisation mechanisms and several optimisations increase the overall performance of dCUDA.The replacement of MPI based communication by a newly designed InfiniBand Verbs based network manager decreases dCUDA's remote communication latency by a factor of 2 to 3. Device local communication latency was improved by a factor of almost 3 thanks to a redesign of the dCUDA notifications system.Performance benchmarks demonstrate that the optimised framework using the InfiniBand network manager outperforms MPI based dCUDA implementations.Depending on communication patterns, it shows about 40% to 90% improved performance to traditional state of the art GPU cluster programming approaches.i