A New Criterion Using Information Gain for Action Selection Strategy in Reinforcement Learning

Kazunori Iwata, Kazushi Ikeda, Hideaki Sakai · IEEE Transactions on Neural Networks · 2004

In this paper, we regard the sequence of returns as outputs from a parametric compound source. Utilizing the fact that the coding rate of the source shows the amount of information about the return, we describe l-learning algorithms based on the predictive coding idea for estimating an expected information gain concerning future information and give a convergence proof of the information gain. Using the information gain, we propose the ratio w of return loss to information gain as a new criterion to be used in probabilistic action-selection strategies. In experimental results, we found that our w-based strategy performs well compared with the conventional Q-based strategy.

Read the paper · More papers on PaperTik