Topology-Aware MPI Communication and Scheduling for High Performance Computing Systems

Hari Subramoni · OhioLink ETD Center (Ohio Library and Information Network) · 2013

Most of the traditional High End Computing (HEC) applications and current petascale applications are written using the Message Passing Interface (MPI) programming model.Consequently, MPI communication primitives (both point to point and collectives) are extensively used across various scientific and HEC applications.The large-scale HEC systems on which these applications run, by necessity, are designed with multiple layers of switches with different topologies like fat-trees (with different kinds of over-subscription), meshes, torus, etc.Hence, the performance of an MPI library, and in turn the applications, is heavily dependent upon how the MPI library has been designed and optimized to take the system architecture (processor, memory, network interface, and network topology) into account.In addition, parallel jobs are typically submitted to such systems through schedulers (such as PBS and SLURM).Currently, most schedulers do not have the intelligence to allocate compute nodes to MPI tasks based on the underlying topology of the system and the communication requirements of the applications.Thus, the performance and scalability of a parallel application can suffer (even using the best MPI library) if topology-aware scheduling is not employed.Moreover, the placement of logical MPI ranks on a supercomputing system can significantly affect overall application performance.A naive task assignment can result in poor locality of communication.Thus, it is important to design optimal mapping schemes with topology information to improve the overall application performance and scalability.It is also critical for users of High Performance Computing ii (HPC) installations to clearly understand the impact IB network topology can have on the performance of HPC applications.However, no currently existing tool allows users of such large scale clusters to analyze and to visualize the communication pattern of their MPI based HPC applications in a network topology-aware manner.This work addresses several of these critical issues in the MVAPICH2 communication library.It exposes the network topology information to MPI applications, job schedulers and users through an efficient and flexible topology management API called the Topology Information Interface.It proposes the design of network topology aware communication schemes for multiple collective (Scatter, Gather, Broadcast, Alltoall) operations.It proposes a communication model to analyze the communication costs involved in collective operations on large scale supercomputing systems and uses it to model the costs involved in the Gather and Scatter collective operations.It studies the various alternatives available to a middleware designer for designing network topology and speed aware broadcast algorithms.The thesis also studies the challenges involved in designing network topologyaware algorithms for network intensive Alltoall collective operation and propose multiple schemes to design a network-topology-aware Alltoall primitive.It carefully analyzes the performance benefits of these schemes and evaluate their trade-offs across varying application characteristics.The thesis also proposes the design of a network topology-aware MPI communication library capable of leveraging the topology information to improve communication performance of point-to-point operations through network topology-aware placement of processes.It describes the design of network topology-aware plugin for the SLURM job scheduler to make scheduling decisions and allocate compute resources as well as place processes in a network topology-aware manner.It also describes the design

Read the paper · More papers on PaperTik