Characterizing node orderings for improved performance

Carl Albing · 2015

High Performance Computing (HPC) job performance can vary because of location in the HPC system interconnect. Contiguous and compact allocations of compute nodes for parallel jobs is ideal but infeasible once other jobs have been placed. Reasonable performance of parallel jobs has been achieved with non-contiguous job placement in 3D-torus HPC systems using allocation strategies based on an ordered, one-dimensional sequence of nodes. Ordering of this list is an inexpensive way to incorporate topological information into the placement decision. With several orderings from which to choose - and others that could be created - what is the basis for choosing one ordering over others? Can the choice be made with out expensive, time-consuming benchmarks?

Read the paper · More papers on PaperTik