Scaling scientific applications on clusters of hybrid multicore/GPU nodes
Lingyuan Wang, Miaoqing Huang, Vikram K. Narayana, Tarek El‐Ghazawi · 2011
Rapid advances in the performance and programmability of graphics accelerators have made GPU computing a compelling solution for a wide variety of application domains. However, the increased complexity as a result of architectural heterogeneity and imbalances in hardware resources poses significant programming challenges in harnessing the performance advantages of GPU accelerated parallel systems. Moreover, the speedup derived from GPU often gets offset by longer communication latencies and inefficient task scheduling. To achieve the best possible performance, a suitable parallel programming model is therefore essential.