Reinforcement Learning for Routing in Communication Networks

Walter H. Andrag · SUNScholar (Stellenbosch University) · 2003

Routing policies for packet-switched communication networks must be able to adapt to changing traffic patterns and topologies.We study the feasibility of implementing an adaptive routing policy using the Q-Learning algorithm which learns sequences of actions from delayed rewards.The Q-Routing algorithm adapts a network's routing policy based on local information alone and converges toward an optimal solution.We demonstrate that Q-Routing is a viable alternative to other adaptive routing methods such as Bellman-Ford.We also study variations of Q-Routing designed to better explore possible routes and to take into consideration limited buffer size and optimize multiple objectives.

Read the paper · More papers on PaperTik