Towards Predictable Energy and Runtime Estimation for Vertically Scaled Edge AI Inference

Uwe Gropengießer, Thomas Reuter, Dominik Schön, Osama Abboud, Xun Xiao, Max Mühlhäuser · 2026

Pervasive edge inference must operate under tight energy and latency budgets while being evaluated against task-dependent result quality. This paper investigates whether per-request latency and energy under vertical CPU scaling can be reliably estimated from lightweight features available prior to execution. To this end, we collect per-request labels for execution time and container-attributed energy in a Kepler-instrumented Kubernetes setup and link them to model variant, CPU allocation, and observable input characteristics. We compute Quality of Result (QoR) offline per model variant on labeled reference data to compare cost profiles and predictive behavior along a reproducible quality axis. Building on this, we benchmark multiple regression methods and compare their performance across configurations. The results show that a substantial fraction of the variation in latency and energy is explainable and predictable from pre-execution features, and identify configurations with increased errors that matter for conservative planning.

Read the paper · More papers on PaperTik