IPDPS 2023 PhD Forum Welcome and Abstracts
2023
Recent advances in multi and many-core processors have led to significant improvements in the performance of scientific computing applications.However, the addition of a large number of complex cores have also increased the overall power consumption, and power has become a firstorder design constraint in modern processors.While we can limit power consumption by simply applying software-based power constraints, applying them blindly will lead to non-trivial performance degradation.To address the challenge of improving the performance, power, and energy efficiency of scientific applications on modern multi-core processors, we propose a novel Graph Neural Network based auto-tuning approach that simultaneously optimizes for runtime performance and energy efficiency by minimizing the energy-delay product.The key idea behind this approach lies in modeling parallel code regions as flow-aware code graphs to capture both semantic and structural code features.In addition, we also use a small set of performance counters to enable our models to capture runtime features of a code kernel.We evaluate our approach on 30 benchmarks and proxy-/mini-applications with 68 OpenMP code regions.Our approach identifies OpenMP configurations for energy-delay product, that lead to performance improvement of 21% and energy reduction of 29% over the default OpenMP configuration at Thermal Design Power for a 32-core Skylake processor.