Estimating Passive Dynamics Distributions and State Costs in Linearly Solvable Markov Decision Processes during Z Learning Execution
Mauricio Alexandre Parente Burdelis, Kazushi Ikeda · SICE Journal of Control Measurement and System Integration · 2014
Although the framework of linearly solvable Markov decision processes (LMDPs) reduces the computational complexity in reinforcement learning, it requires the knowledge of the state-transition probability in the absence of control or passive dynamics. The passive dynamics can be estimated by a temporal difference method called Z learning if the environment obeys the passive dynamics. However, it leads to a slow convergence of learning since no control is allowed during learning. This paper proposed a method to estimate the passive dynamics using Z learning under a different state-transition probability from the passive dynamics. The proposed method requires only the knowledge on what states can be visited from each possible state, and estimates the state-transition probability as well as the immediate cost of the states from the constraints they should satisfy. The computer experiments showed that the proposed method remains more efficient than Q learning with successful estimation of the passive dynamics and state costs and has a comparable convergence speed with the traditional Z learning.