Demonstration of a primitive dissipating-predictive-state planning actor-critic in a game of cat and mouse

Justin Schlauwitz, Petr Musı́lek · 2017

This article covers the comparison of a primitive Dissipating-Predictive-State Planning Actor-Critic learning approach with SARSA(λ) (State-Action-Reward-State-Action) on the highly non-stationary competitive Cat and Mouse problem. The primary objective of the new algorithm was to minimize the number of constants that must be optimized before application while maintaining the performance found in other basic learning algorithms such as SARSA. The new approach was able to perform to a satisfactory degree despite the applied algorithm and parameters not being able to demonstrate better overall performance when compared to SARSA. However, from the analysis of this contribution, we were able to gather several important insights to the operations of these two algorithms. Additionally, it must be admitted that there are still many areas of the new algorithm that can be optimized. It would also be worth testing it on a few toy problems, such as grid-world, that have a greater degree of stationarity.

Read the paper · More papers on PaperTik