Confidence based dual reinforcement Q-routing: an adaptive online network routing algorithm
Shailesh Kumar, Risto Mukkulainen · 1999
This paper describes and evaluates the Confidence-based Dual Reinforcement QRouting algorithm (CDRQ-Routing) for adaptive packet routing in communication networks. CDRQ-Routing is based on an application of the Q-learning framework to network routing, as first proposed by Littman and Boyan (1993). The main contribution of CDRQ-routing is an increased quantity and an improved quality of exploration. Compared to Q-Routing, the state-of-the-art adaptive Bellman-Ford Routing algorithm, and the non-adaptive shortest path method, CDRQ-Routing learns superior policies significantly faster. Moreover, the overhead due to exploration is shown to be insignificant compared to the improvements achieved, which makes CDRQ-Routing a practical method for real communication networks. 1 Introduction In a communication network information is transferred from one node to another as data packets [ Tanenbaum, 1989 ] . The process of sending a packet P (s; d) from its source node s to its destination node d ...