Complex-Valued-based Learning Classifier System for POMDP Environments

Keiki Takadama, Daichi Yamazaki, Masaya Nakata, Hiroyuki Satō · 2019

This paper proposes Complex-Valued-based Learning Classifier System (CVLCS) that can learn an appropriate policy for the POMDP environments by extending Complex-Valued Reinforcement Learning (CVRL). Concretely, CVLCS explores the optimal policy by not only evolving classifiers but also updating Q-values (i.e., strength) of evolved ones, while CVRL explores the optimal policy by only updating Q-values of the state-action pairs prepared beforehand. To investigate the effectiveness of CVLCS, this paper applies it to various types of the POMDP environments. The experimental results have revealed that (1) CVLCS can derive the good performance which is close to the optimal one and shows such a performance faster than the conventional methods (i.e., Q-Learning as one of CVRL and ZCSM as one of LCS) and (2) CVLCS can stably derive the good performance even in the difficult environments where the conventional methods fail to derive good performance.

Read the paper · More papers on PaperTik