Finite Sample Analysis of Minmax Variant of Offline Reinforcement Learning for General MDPs

Jayanth Reddy Regatti, Abhishek Gupta · IEEE Open Journal of Control Systems · 2022

In this work, we analyze the finite sample complexity bounds for offline reinforcement learning with general state, general function space and state-dependent action sets. The algorithm analyzed does not require the knowledge of the data-collection policy as compared to earlier works. We show that one can compute an$\epsilon$-optimal Q function (state-action value function) using$O(1/\epsilon ^{4})$i.i.d. samples of state-action-reward-next state tuples.

Read the paper · More papers on PaperTik