Improved Parallel Application Performance and Makespan by Colocation and Topology-aware Process Mapping

Ioannis Vardas, Sascha Hunold, Philippe Swartvagher, Jesper Larsson Träff · 2024

In modern, deeply hierarchical HPC systems shared resource congestion can hinder the efficient use of many cores by parallel applications and degrade performance. Such congestion is often caused when parallel processes within an application that execute similar operations share the same resources. Previous research suggests using fewer cores with better process-to-core mapping can improve applications’ performance but leaves many cores unused. To utilize these cores, we colocate additional applications and map them using a topology-aware process-to-core, application-agnostic mapping algorithm. We show that these mappings significantly impact memory bandwidth and communication latency. We evaluate our approach using eight parallel applications on an HPC system with 128-core nodes, demonstrating the performance effects of mappings combined with colocation. Our goal is to determine whether colocation with topology-aware mapping is a viable alternative to typical exclusive node allocation. Our results show makespan improvements of 2.4x over exclusive allocation in an HPC system, demonstrating the potential benefits of colocation with optimized mappings.

Read the paper · More papers on PaperTik