Q-learning policies for a single agent foraging tasks
M. Yogeswaran, S. G. Ponnambalam · International Symposium on Mechatronics and its Applications · 2010
Policies play an important role in balancing the trade-off between exploration and exploitation problem in q-learning. Pure exploration degrades the performance of the q-learning but increases the flexibility to adapt in a dynamic environment. On the other hand pure exploitation drives the learning process to locally optimal solutions. In this paper, a single agent foraging task has been modeled incorporating the available policies reported in the open literature to address the exploration and exploitation issues. Policies namely greedy, e-greedy, Boltzmann distribution, Simulated An-nealing(SA)algorithm and random search are used to study their performances in the foraging task and the results are presented.