Distributional-Utility Actor-Critic for Network Slice Performance Guarantee
Jingdi Chen, Tian Lan, Nakjung Choi · 2023
Optimizing distributional utilities (such as mitigating performance tails and maximizing risk-aware objectives) is crucial for online network slice management to meet the diverse requirements of different services and applications. While Reinforcement Learning (RL) has been successfully applied to autonomous online decision-making in many network slice management problems, existing solutions often focus on maximizing the expected cumulative reward or are limited to specific distributional utilities. This paper proposes a new RL algorithm for general Distributional Utilities Optimization (DUO) in an actor-critic framework for online network slice management. In particular, we derive a DUO Temporal Difference Learning algorithm for updating distributional utilities in the critic through stochastic gradient descent. It is proven that the Distributional Optimal Bellman Operator for distributional utilities is a γ-contraction and thus is guaranteed to converge. In addition, we parameterize the policy by another neural network and prove a revised policy gradient theorem for distributional utilities, which shows that the derived policy update converges to at least a stationary point of the DUO problem. Our proposed algorithm works with arbitrary smooth utility functions on the return distributions, making it suitable for optimizing various network slice performance objectives in an online setting. Our solution is implemented and validated by building a hybrid trace-driven network simulator, which was built using an open-source O-RAN dataset, along with data collected from a 5G O-RAN testbed. Results demonstrate a significant improvement over heuristic and RL baselines.