Payo size variation problem in simple reinforcement learning algorithms 1
Michal Kvasnička · 2013
This paper shows that the speed of the reinforcement learning depends on the size of the payoffs, at least when all payoffs are positive. When the speed of learning is too fast, the agents tend to learn to play the actions which they randomly chosen in the first rounds of the learning process. The compositions of the agents’ strategies then on the aggregate level resembles the initial individual agent’s mixed strategy. This may create artificial effects in the simulations where the size of payoffs depend on the model treatments because the speed of learning cannot be tuned in.