SLIP: A Sophisticated Learner for Instance-based Policy using Hybrid GA

Chikao Tsuchiya, Yusuke Shiokawa, Kokolo Ikeda, Jun Sakuma, Isao Ono, Shigenobu Kobayashi · Transactions of the Society of Instrument and Control Engineers · 2006

Reinforcement learning is a useful tool for complex control problems that cannot be modeled mathematically nor solved theoretically. However, a traditional value function approach such as Q-learning includes the difficulty of combinatorial explosion. Direct policy search, (DPS) is an alternative approach that represents a policy using some model and searches a parameter space directly for an optimum by optimization techniques such as genetic algorithms (GA). Instance-based policy (IBP) is a policy representation model of DPS. IBP represents a policy using a set of instances that are pairs of state and action. This paper presents a hybrid GA to optimize efficiently a set of instances with continuous state and continuous action, given an episodic task. The hybrid GA is composed of a combinatorial GA with BDX (Binomial Distribution Crossover) and a real-coded GA with INDX (Instance-wise Normal Distribution Crossover). The proposed method named SLIP (Sophisticated Learner for Instance-based Policy) was applied to a cat twist problem and a parallel-type double inverted pendulum problem.The results of experiments show the effectiveness and usefulness of SLIP.

Read the paper · More papers on PaperTik