Optimizing Auto-tuning of OpenMP Offload kernels for performance and power
Nafis Mustakin, Daniel Lin-Kit Wong · 2025
Auto-tuning frameworks for GPGPU kernels typically focus on minimizing runtime by adjusting easily-accessible parameters such as kernel size.However, they often converge to suboptimal configurations due to limited visibility into the broader performance landscape.In this work, we augment the auto-tuning process by incorporating energy-aware profiling data -including memory load/store intensity, warp occupancy, and branch efficiency -into the optimization loop.We target energy (E), energy-delay product (EDP), and energy-delay-squared product (ED2P) as tuning objectives, and demonstrate that our context-enriched framework consistently identifies kernel configurations that are both faster and more energy-efficient.Our evaluation shows that different kernels exhibit varying sensitivity to these objectives, and that profiling metrics can serve as effective predictors of which tuning targets are likely to yield performance gains.By integrating these metrics as optimization objectives, we enable more informed and robust kernel tuning across diverse workloads.