Reinforcement Learning with Sparse Bellman Error Extrapolation for Infinite-Horizon Approximate Optimal Regulation

Max L. Greene, Patryk Deptula, Scott Nivison, Warren E. Dixon · 2019

This paper provides an approximate online adaptive solution to the infinite-horizon optimal control problem for control-affine continuous-time nonlinear systems. The state-space is segmented into a user-defined number of segments. Off-policy trajectories are selected over each segment to facilitate learning of the value function weight estimates. Sparse neural networks enable a framework for switching and state space segmentation as well as computational benefits due to the small number of neurons that are active. At each sparse segment, the off-policy trajectories are used to extrapolate the Bellman error (BE) across their respective segments to provide an optimal policy for each segment. Over each segment, the extrapolated BEs are used in the value function weight update laws. Because at each segment a different set of extrapolated BEs is used in the update laws, discontinuities occur in the weight update laws. A Lyapunov-like stability analysis is included which proves boundedness of the overall system in the presence of discontinuities.

Read the paper · More papers on PaperTik