Reinforcement learning in the environment where optimal action value function is partly discontinuous
Shingo Shibusawa, Takeshi Shibuya · 2016
The application of reinforcement learning to robot control which has continuous state and action requires approximation of action value function. Radial basis function networks (RBFN) is one of the methods for function approximation. However, it makes an agent, a learning subject, to select less valuable actions near the states where optimal action value function is discontinuous. To solve this problem, this paper proposes a method to divide the states into two regions and to use the different weight vectors in each region. States are divided by the straight line which expresses the distribution of states where optimal action value function is discontinuous. The parameter of the line is estimated by the histories of states, actions, and TD errors. The proposed method enables agent to select valuable actions near those states.