Designing Internal Reward of Reinforcement Learning Agents in Multi-Step Dilemma Problem
Y. Ichikawa, Keiki Takadama · Journal of Advanced Computational Intelligence and Intelligent Informatics · 2013
This paper proposes the reinforcement learning agent that estimates internal rewards using external rewards in order to avoid conflict in multi-step dilemma problem. Intensive simulation results have revealed that the agent succeeds in avoiding local convergence and obtains a behavior policy for reaching a higher reward by updating the Q-value using the value that is subtracted the average reward from an external reward.