Predicting Power and Timing of Large-Scale Distributed Applications on Highly Heterogeneous Platforms
Walter, Jörg · Carl von Ossiezky University of Oldenburg · 2020
Applications in the high-performance computing (HPC) domain are often designed to run on cluster-like distributed platforms with hundreds of nodes. Due to the size of both – applications and platforms – predictions of application run time and energy usage is challenging. Furthermore, most HPC prediction methodologies only address timing, because energy predictions used to have little relevance for HPC application design. I propose a new approach to this challenge based on a simulation technique well known in the embedded computing domain. In order to apply such a methodology at HPC scale, I cannot execute actual applications during simulation. I use abstract simulation based on the Task Graph model of computation, which is popular in HPC and which has properties similar to the synchronous dataflow model that is popular in the embedded domain.