Reinforcement learning with particles for instant optimality
Tsuyoshi Beppu, Akira Notsu, Katsuhiro Honda, Hidetomo Ichihashi · 2012
In this paper, we propose a new Actor-Critic method in the agent environment and action space based on the normal Actor-Critic method and PSO. In the algorithm, particles are expressed as cluster center of some states or actions, and explore through the space in order to get an appropriate divided space. The purposes of this study are learning efficiency improvement and heuristic space segmentation. In our method, particles move in the space during the agent's learning process. Appropriate segmentation can minimize the learning time and enables us to recognize the evolutionary process. Thus, this method is also designed for humanlike decisions in the learning process. The simulation results indicate that our method shows some clusters in the action and state space. Space segmentation, such as group formation, language systems and culture, will be revealed by multi-agent social simulation with our method.