Cycle-Stealing in Load-Imbalanced HPC Applications
Po Hao Chen, Akshaya Bali, Shining Yang, Pouya Haghi, Carlton Knox, B Li, Amr Akmal Abouelmagd, Anthony Skjellum, Martin Herbordt · 2024
It is practical to steal cycles when Message Passing Interface (MPI) programs are load imbalanced, either from natural algorithmic load, because of collective communication and/or process skew, and/or because of variable message loads and needs for message progress within the MPI implementation. We introduce TimeLord, a runtime library to provide fungibility for compute-cycles without significantly degrading the nominal performance of the parallel application. We propose three means to exploit wasted cycles in LAMMPS, miniFE, and miniAMR. The results show, on average, 40% of the runtime can be used to execute extra computations and that runtime overheads are less than 4%. We envisage that these extra computations be related to the target application; here, for simplicity and proof-of-concept, we assume arbitrary commodity applications. Additionally, TimeLord uncovers inefficiencies in the MPI progress engine and improves baseline performance. We explore some fundamental issues of load imbalance and its interaction with system software, including the predictability of extra time and the relative benefits of prediction versus preemption. TimeLord requires no changes to the MPI application, MPI implementation, or other system components.