Latency-Aware Placement of Microservices in the Cloud-to-Edge Continuum via Resource Scaling
Alberto Bertoncini, Alberto Ceselli, Christian Quadri · 2025
Latency-sensitive applications, such as autonomous driving in smart cities and smart industries, require a networking and computing infrastructure to support their operations. Cloud-to-edge continuum represents a promising architecture to provide computational capability close to edge devices. However, deploying latency-sensitive applications in the continuum is challenging due to the heterogeneity and the geographical distribution of the computing nodes. In this paper, we address the deployment problem in a tele-operated autonomous driving scenario, formulating the orchestration task as a Virtual Network Function Placement Problem (VNFPP) with multi-tier performance levels, enabling vertical scaling of computational resources per microservice. Our MILP model, MORAL, minimizes node centrality-based deployment costs while satisfying resource and end-to-end latency constraints. We tested our approach through extensive simulations on realistic network topologies and synthetic applications, showing that the proposed model improves deployment feasibility, latency compliance, and resource efficiency compared to single performance tier versions and baseline strategies.