Transfer Learning for Constrained Stochastic Control Using Adjustable Benders Cuts

Joseph Warrington · IEEE Control Systems Letters · 2021

Stochastic optimal control problems are generally difficult to solve to optimality, and it is desirable to be able to transfer a high-quality solution learned for one such problem to a new, related problem at low computational cost. This is useful in cases such as a known change in dynamics, e.g., when a vehicle's payload increases. We describe a way of doing this for discrete time, infinite horizon problems on continuous spaces. The solutions take the form of Q-function approximations, consisting of “Benders cuts” that lower-bound the optimal Q-function. This form induces an easy to evaluate control policy. We describe how to adjust a set of existing cuts at low computational cost while preserving the lower bounding property for the new problem. This transfers prior training effort to the new problem far more cheaply than starting from scratch. We use random problem instances to demonstrate that our transfer procedure yields near-optimal results on the new model, outperforming a clipped linear-quadratic regulator (LQR) on average. The Benders approach, while model-based, is strongly related to model-free reinforcement learning, in that it uses sampled experience evaluating the so-called Bellman error to learn the value function.

Read the paper · More papers on PaperTik