Transition Entropy in Partially Observable Markov Decision Processes.
Francisco S. Melo, Isabel Ribeiro · 2006
This paper proposes a new heuristic algorithm suitable for real-time applications using partially observable Markov decision processes (POMDP). The algorithm is based in a reward shaping strategy which includes entropy information in the reward structure of a fully observable Markov decision process (MDP). This strategy, as illustrated by the presented results, exhibits near-optimal performance in all examples tested.