WMat algorithm based on Q-Learning algorithm in taxi-v2 game
Fatemeh Esmaeily, Mohammad Reza Keyvanpour · 2020
Reinforcement learning is a framework in which an agent aims to optimize the sum of the rewards it receives from the environment around based on a specific policy. Although many researches have been done on this type of learning, it has not been specifically addressed to improve Q-table values and to increase the amount of reward received by the agent. One of the notable issues in this regard is the improvement of the results of the q-learning algorithm functions. Hence in order to improve the amount of received reward, a weighting matrix is proposed and defined for the set of possible actions of the agent. In this regard our method is capable of reducing the likelihood of behaviours consideration by the agent which leads to a notable decrease in the time of action choosing and as a result achieves improvement in the amount of reward and produces acceptable results in this criteria.