Reinforcement Learning Solution to a Benchmark Time-Optimal Control Problem
D. Duane Dier, Sanjay S. Joshi · AIAA Guidance, Navigation, and Control Conference and Exhibit · 2007
Reinforcement learning methods originated with reward-punishment studies in psychology, and were then extended to machine learning algorithms in computer science. The advantage of reinforcement learning methods is that they do not require any knowledge of a system’s dynamics, and use experience gained from interaction with the actual system (or simulation thereof) to obtain control solutions. In fact, strict RL and its derivatives have inspired several new methodologies, which have begun to be used on complex aerospace systems. In this paper, we apply traditional RL to a well-known simply-posed minimum time optimal control problem using the Sarsa(λ) reinforcement learning method. It is wellknown by control researchers that the true analytic optimal solution is a “bang-bang” solution. In fact, analytical proof of optimality for Sarsa(λ) has yet to be achieved for either discrete state or continuous state optimal control problems (though it is an active area of research). The current study showed that Sarsa(λ) did produce nearly-optimal “bang-bang” results for the given benchmark problem—without any explicit a-priori knowledge of the system dynamics. However, generalization of the numerical solution from a single initial condition to other initial conditions was not immediate.