Fear based Intrinsic Reward as a Barrier Function for Continuous Reinforcement Learning

Rodney Sanchez, Ferat Sahin, Jamison R. Heard · 2024

Current reinforcement learning (RL) methods must explicitly sample states to learn about their value. Redundantly sampling these states creates dangerous situations when RL agents are deployed in real-life environments. Furthermore, since the agent only receives a reward when entering the environment, the agent must sample the state multiple times to modify the agent’s policy to avoid dangerous states. Humans, specifically young children, primarily overcome this need for explicit and redundant sampling using two key strategies: fear and vicarious conditioning. Our method utilizes fear and vicarious conditioning to create a pseudo-barrier function that discourages the agent from sampling the dangerous state. Using memory augmented neural network (MANN) similarity calculations, we can see how similar the agent’s current state is to the “phobia” by creating a dense reward field that serves as a pseudo-barrier function. The MANN was trained in the MiniGrid Simple environment, while the agent was tested in the LAVAGAP and Dynamic Obstacle environments. Our results show that the MANN can produce a dense reward gradient that transfers to different Minigrid environments. Our method also shows that this “phobia” can discourage the agent from visiting certain states.

Read the paper · More papers on PaperTik