Ex〈α〉: An effective algorithm for continuous actions Reinforcement Learning problems
Jose Antonio Martin H, Javier de Lope · 2009
In this paper the Ex(α) Reinforcement Learning algorithm is presented. This algorithm is designed to deal with problems where the use of continuous actions have clear advantages over the use of fine grained discrete actions. This new algorithm is derived from a baseline discrete actions algorithm implemented within a kind of κ-nearest neighbors approach in order to produce a probabilistic representation of the input signal to construct robust state descriptions based on a collection (knn) of receptive field units and a probability distribution vector p(knn) over the knn collection. The baseline continuous-space-discrete-actions kNN-TD(λ) algorithm introduces probability traces as the natural adaptation of eligibility traces in the probabilistic context. Later the Ex(α)(κ) algorithm is described as an extension of the baseline algorithms. Finally experimental results are presented for two (not easy) problems such as the Cart-Pole and Helicopter Hovering.