A solving method for MDPs by minimizing variational free energy
Junichiro Yoshimoto, Shin Ishii · 2005
We propose a novel approach to acquire the optimal policy for a continuous Markov decision process. Based on an analogy from statistical mechanics, we introduce a variational free energy over a policy. A good policy can be obtained by minimizing the variational free energy. According to our approach, the optimal policy in linear quadratic regulator problems can be obtained by using Kalman filtering and smoothing techniques. Even in non-linear problems, a semi-optimal policy can be obtained by Monte Carlo technique with a Gaussian process method.