Recurrent neural networks for reinforcement learning: architecture, learning algorithms and internal representation
Ahmet Onat, Hajime Kita, Yoshikazu Nishikawa · 2002
Reinforcement learning is a learning scheme for an autonomous agent that allows the agent to find the optimal policy of taking actions which maximize a scalar reinforcement signal in unknown environments. If the agent has access to the whole state of the environment, a reactive policy which maps the sensory input to the action is sufficient. However, if the state of the environment is partially observable, special methods for creating a dynamic policy that utilizes the past observations are necessary. To overcome this problem, the authors have proposed a method using recurrent neural networks with Q-learning, as a learning agent. The paper compares several types of network architecture and learning algorithms for this method through computer simulation. Further, the internal representation in the trained networks is examined using a clustering technique. It shows that the representation of the environmental state is developed well in the networks.