Topology Mapping of Parallel Applications onto Random Allocations
Yao Hu · 2019
In high-performance computing (HPC) systems, one application usually has many parallel tasks running on multiple compute nodes. The execution time depends on the communication latency and the network contention between the parallel tasks especially when these tasks intensively communicate with each other. Traditionally, the tasks are mapped onto a regular topology with nearby compute nodes to reduce the network distances. In this case, fragmentation of unused compute nodes cannot be assigned for a newly incoming job, because it is assumed to largely harm the communication abilities between non-adjacent nodes. However, the communication overhead between communicating tasks is also determined by the application's communication pattern, which may not suit to a specified regular topology. For the purpose of improving job scheduling abilities, we investigate the job mapping on random topologies so that a newly incoming job can be immediately dispatched as long as there are enough available nodes. Evaluation results show that, for a large compound workload of NAS Parallel Benchmarks (NPB) applications, the random job mapping can reduce up to 64% of makespan and up to 80% of turnaround time when compared with the regular job mapping. Overall, the random topology embedding in random topologies indicates substantial room for improvement of job scheduling performance.