Multi-order Rule Accumulation for an Agent Control Problem in Non-Markov Environments
Lutao Wang, Shingo Mabu, Kotaro Hirasawa · IEEJ Transactions on Electronics Information and Systems · 2012
Multi-agent control in non-Markov environments is difficult because the environment information is partially observable. Agents suffer from the perceptual aliasing problem and couldn't take proper actions. In order to solve this problem, this paper proposes a rule-based model named “multi-order rule accumulation” to guide agent's actions in non-Markov environments. The advantages are, firstly, each multi-order rule memorizes the past environment information and agent's actions, which serves as the additional information to distinguish the aliasing situations, secondly, multi-order rules are very general, so that they are competent for guiding agents' actions in Partial Observable Markov Decision Process (POMDP), thirdly, multi-order rules are accumulated throughout the generations, which could cover many situations experienced in different generations. This also helps agents to take proper actions. Simulations on the tile-world problem prove that this rule-based model outperforms the conventional methods and the previous research.