A TWO-DIMENSIONAL CONTINUOUS ACTION LEARNING AUTOMATON
K. Spurgeon, Qinghua Wu, Z. Richardson, J. Patrick Fitch · 2004
This paper presents an expansion of the Continuous Action Reinforcement Learning Automaton to incorporate two-dimensional (2D) actions. The most significant change to the original CARLA methodology is the introduction of a matrix J that is used to store the last known reward value for each action. Inference rules based on comparisons between the values held in J, the current reward value and the probability distribution function of the automata are then used to directly guide the non-linear Reward-Penalty scheme. The automaton is tested firstly in an environment where the optimum is subject to noise and secondly tracking a time varying optimal. In both cases the learning algorithm performs well, providing significant reinforcement to the current optimum whilst simultaneously avoiding saturation even after long periods.