The Curse of (Too Much) Choice: Handling Combinatorial Action Spaces in Slice Orchestration Problems Using DQN with Coordinated Branches
Pavlos Doanis, Thrasyvoulos Spyropoulos · 2024
One of the prominent problems in envisioned 6G networks is the truly dynamic placement of multiple virtual network function chains on top of the physical network infrastructure. Reinforcement Learning based schemes have been recently explored for such problems. Yet these have to deal with astronomically high state and action spaces in this context. Using a standard Deep Q-Network (DQN) is a common way to effectively deal with state complexity. While the use of independent DQN (iDQN) agents could be further used to mitigate action space complexity, such schemes often suffer from instability and sample (in)efficiency, and their theoretical performance is hard to assess. To this end we propose a DQN-based scheme that uses a recent Deep Neural Network architecture, with a different branch responsible for the placement of each virtual network function (again reducing action space complexity), yet with (implicit) coordination among branches, via shared layers (hence avoiding iDQN shortcomings). Using a real traffic dataset, we (i) theoretically ground the proposed scheme by comparing it with an optimal online algorithm for a stateless experts environment; (ii) we demonstrate a 41% cost improvement compared the existing state-of-the-art multi-agent DQN approach (independent agents).