An Intrinsic Motivation Based Artificial Goal Generation in On-Policy Continuous Control
Baturay Sağlam, Furkan B. Mutlu, Kaan Gönç, Onat Dalmaz, Süleyman S. Kozat · 2022 30th Signal Processing and Communications Applications Conference (SIU) · 2022
This work adapts the existing theories on animal motivational systems into the reinforcement learning (RL) paradigm to constitute a directed exploration strategy in on-policy continuous control. We introduce a novel and scalable artificial bonus reward rule that encourages agents to visit useful state spaces. By unifying the intrinsic incentives in the reinforcement learning paradigm under the introduced deterministic reward rule, our method forces the value function to learn the values of unseen or less-known states and prevent premature behavior before sufficiently learning the environment. The simulation results show that the proposed algorithm considerably improves the state-of-the-art on-policy methods and improves the inherent entropy-based exploration.