Distributed reinforcement learning for a traffic engineering application
Mark D. Pendrith · 2000
In this paper, we report on novel reinforcement learning techniques applied to a real-world application.The problem domain, a traffic engineering application, is formulated as a distributed reinforcement learning problem, where the returns of many agents are simultaneously updating a single shared policy.Learning occurs off-line in a traffic simulator, which allows us to retrieve and exploit good transient policies even in the presence of instabilities in the learning.We introduce two new algorithms developed for this situation, one which is a value function based, and one that employs a direct policy evaluation approach.While the latter is theoretically better motivated in several ways than the former, we find both perform comparably well in this domain and for the formulation we use. 1