Distributed Task-Based Runtime Systems - Current State and Micro-Benchmark Performance
Reazul Hoque, Pavel Shamis · 2018
High Performance Computing Systems are moving heavily towards many-core processors with a deep hierarchy of memory. Accelerators like GPUs are widely being used for general purpose computing and processor architectures are becoming increasingly complex to accommodate performance boost. This trend towards complex heterogeneous architecture makes the job of scientific application developers difficult in terms of performance, portability and productivity. With memory being distributed, this challenge becomes even more complex. Programming many-core shared memory systems are most widely accomplished using OpenMP, while MPI is used to manage the communications in a distributed system. Even though MPI provides a rich set of features, it is too explicit making users responsible for overlapping communication and computation. Task-based runtime systems have emerged as a solution to this challenge of programming these modern complex systems. This study surveys the landscape of task-based runtime systems that support distributed memory and presents a set of benchmark for evaluating and understanding runtime-performance and overheads of these systems.