Inference-Time-Driven Autoscaling for Inference Workloads: A Comparative Study of Latency-Variant Models in Kubernetes
Josephine Eskaline Joyce, Shoney Sebastian · Technologies · 2026
Kubernetes Horizontal Pod Autoscaler (HPA) primarily relies on resource-based metrics, such as CPU utilization, which are poorly suited to capturing the latency variability of AI inference workloads. In this paper, we propose a custom-metric-driven autoscaling approach that leverages inference latency histograms as first-class scaling signals for Kubernetes HPA. The proposed framework integrates a Prometheus Operator (PO)-based observability stack with the Prometheus Adapter to expose and aggregate per-pod inference latency metrics, enabling workload-aware scaling decisions. We evaluate the approach using four mid-scale transformer-based inference services, comprising two reasoning-like and two latency-stable workloads, under high-concurrency conditions. The experiments analyze latency variation, tail behavior, and replica dynamics across multiple autoscaling policies, including variations in scale-up aggressiveness (3 pods/30 s, 3 pods/60 s, 6 pods/60 s), inference-time thresholds, and stabilization windows. Compared to CPU-based autoscaling, inference-driven policies reduce mean response time by 18–27% for reasoning-like workloads and 12–20% for stable workloads. The results show that latency-variable workloads exhibit wider tails and higher variance, indicating the need for moderately aggressive scale-up strategies to avoid long-lasting latency spikes. Overall, the findings show that inference-latency-driven custom metrics significantly improve autoscaling efficiency and stability for transformer-based inference workloads in cloud-native environments.