Optimising MPI tree-based communication for NUMA architectures
Christer Karlsson, Zizhong Chen · International Journal of Autonomous and Adaptive Communications Systems · 2015
Today's computer clusters are often composed of many multi-core processors that are networked together. With this architecture communication between cores on different nodes is often on a magnitude slower than those between cores on the same node. Cores on the same processor communicate faster than cores on different processors on the same node. Most MPI implementations assume a homogeneous network. In this paper, we treat a multi-core node as a heterogeneous unit and optimise MPI scatter/gather communications by scheduling using topology information. We demonstrate that a previous heuristics for heterogeneous clusters do improve the performance, but might not produce optimal results on multi-core node for communications. Our solution modifies the fastest edge first heuristic by accounting for how many messages can be sent in parallel without impeding the bandwidth. We are able to achieve 20% to 30% performance gains over the MPI scatter/gather implementation on homogeneous, multi-core nodes.