Fast Learning Automata

Junqi Zhang, MengChu Zhou · 2023

An update scheme of a state probability vector of actions is critical for a learning automaton (LA). The most popular one is the pursuit scheme that pursues the estimated optimal action and penalizes others. This chapter introduces two LAs to accelerate the convergence and computational update of LAs. The first one achieves significantly faster convergence and higher accuracy than the classical pursuit scheme. The other lowers the computational complexity of updating a state probability vector to be independent of the number of actions. The chapter introduces a reverse philosophy opposed to the traditional pursuit scheme and leads to Last-position Elimination-based Learning Automata where the action graded last in terms of the estimated performance is penalized by decreasing its state probability and is eliminated when its state probability is decreased to zero. The discretized pursuit LA is the most popular one among variants of Learning automata.

Read the paper · More papers on PaperTik