Event-learning with a non-Markovian controller
István Szita, Bálint Máté Takács, András Lörincz · 2002
Abstract. Recently a novel reinforcement learning algorithm called event-learning or E-learning was introduced. The algorithm based on events, which are defined as ordered pairs of states. In this setting, the agent optimizes the selection of desired sub-goals by a traditional value-policy function iteration, and utilizes a separated algorithm called the controller to achieve these goals. The advantage of event-learning lies in its potential in non-stationary environments, where the near-optimality of the value iteration is guaranteed by the generalized ε-stationary MDP model. Using a particular non-Markovian controller, the SDS controller, an ε-MDP problem arises in E-learning. We illustrate the properties of E-learning augmented by the SDS controller by computer simulations. 1