Relative reward strength algorithms for learning automata

Rahul Simha, Jim Kurose · IEEE Transactions on Systems Man and Cybernetics · 1989

A novel class of action probability update algorithms for learning automata that use the relative reward strengths of responses from the environment is examined. Specifically, update algorithms for S-model automata in which 'recent' environmental responses for each of the actions retained are used. A convergence result is proven and the behavior of these automat is studied by simulation. A major result is that the performance of these algorithms is superior, in several respects, to that of the well-known SL/sub R-1/ update algorithm. Additional results are presented on the variability of performance, the cost of learning and, in the case of static environments, modifications that result in improved convergence.>

Read the paper · More papers on PaperTik