Dual Policy-Based TD-Learning for Model Predictive Control
Chang-Hun Ji, Ho-Bin Choi, Joo-Seong Heo, Ju-Bong Kim, Hyun-Kyo Lim, Youn‐Hee Han · 2023
Temporal difference learning for model predictive control (TD-MPC) is the state-of-the-art in data-driven reinforcement learning-based MPC. Although TD-MPC maintains two different policies: 1) a model predictive path integral-based policy and 2) a parameterized policy, it applies only the first policy action to the environment. In this paper, we propose an extension of TD-MPC, called TD-MPC+, which selects a better policy that generates more practical actions, which exploits the agent's current estimated models greedily to get the most reward. It is noted that TD-MPC+ does not require an additional computational cost of model training than the original TD-MPC. In a comparison study with diverse the DeepMind Control Suite tasks, TD-MPC+ has higher sample efficiency and performance than TD-MPC.