Unifying Plan-Space Value-Based Approximate Dynamic Programming Policies and Open Loop Feedback Control

Lawrence A. M. Bush, Brian Charles Williams, Nicholas Roy · AIAA Infotech@Aerospace Conference · 2009

We examine the complementary strengths of value function based policy learning and guided search. Our work unifies rollout based open-loop feedback control (outlined by Bertsekas 2 ) and plan-space approximate dynamic programming (studied by Boyan 3 ). We exploit the strengths of this unified space by finding a natural order and metric for considering plan changes. Open-loop feedback control builds and executes a sequence of actions (an open-loop plan) at runtime. When the resulting real-world situation deviates from what was expected, the system replans. Reinforcement learning solves the same problem by learning an approximate value function, which is subsequently used to select actions. We typically cannot learn exact value functions for interesting real-world problems. Therefore, we use rollout (value function based limited search) to construct an open-loop plan that improves upon value function based action selection. Rollout is performed in the real-world state-space and actions are added to the plan in time order, the same order as they are exectued in the real-world. On the other hand, planspace search constructs an open-loop plan out of order. For example, a future action may be added to a plan even though our immediate action has not yet been determined. The term plan-space refers to the fact that we are searching though a space of plans. We can likewise guide plan-space search by learning a value function. A plan-space value function informs us of actions that should be added to our plan. Rollout-based open-loop feedback control and plan-space approximate dynamic programming are generally considered orthogonal. 7 However, we have constructed an inclusive framework that incorporates both methods. We present experiments to find natural planning action spaces, which prescribe allowable plan changes that meaningfully reflect their impact on global optimality. We demonstrate our approach on multiple unmanned air vehicle adaptive mission planning.

Read the paper · More papers on PaperTik