Accurate HPC Network Simulations Using Application-Level Approximation

Elkin Cruz-Camacho, Christopher D. Carothers · 2024

To keep up with the times, supercomputers have to evolve as do the applications they enable in science, national security and artificial intelligence. To do so, the best strategy we have is to simulate how changes in their algorithms and architecture design would impact their performance. Yet, simulation at high fidelity, keeping track of all the interlocking parts in detail, requires in itself large computing resources; just a couple of milliseconds of simulated time can take hours to run. Fortunately, we can exploit massively parallel applications tendency to follow an iterative pattern with clearly defined stages. Every iteration takes roughly the same amount of resources including network utilization. We can record the resources utilized during a couple of iterations and from them we can estimate how long it will take to continue the simulation for dozens or hundreds of iterations longer. Determining how long will each iteration take highly depends on the placement, number of iterations and kinds of applications running alongside. We implement a strategy to record the state of the network, train a statistical surrogate model to estimate the time each iteration will take, and switch the simulation into a low-fidelity, surrogate-enabled mode in which we have seen gains close to the number of iterations skipped, i.e, if we skip 80% of iterations we see a speedup of 5 ×.

Read the paper · More papers on PaperTik