Characterization of the absolutely expedient learning algorithms for stochastic automata in a non-discrete space of actions.
Carlos Rivero · The European Symposium on Artificial Neural Networks · 2003
This work presents a learning algorithm to reach the optimum action of an arbitrary set of actions contained in IRm. An initial and arbitrary probability measure on IRm allow us to select an action and the probability is sequentially updated by a stochastic automaton using the response of the environment to the selected action. We prove that the corresponding random sequence of probability measures converges in law to a probability measure degenerate on the optimum action, with probability as close to one as we desire.