Policy Improvement: Between Black-Box Optimization and Episodic Reinforcement Learning

Freek Stulp, Olivier Sigaud · 2013

Abstract: Policy improvement methods seek to optimize the parameters of a policy with respect to a utility function. There are two main approaches to performing this optimization: reinforcement learning (RL) and black-box optimization (BBO). In recent years, benchmark comparisons between RL and BBO have been made, and there have been several attempts to specify which approach works best for which types of problem classes. In this article, we make several contributions to this line of research by: 1) Classifying several RL algorithms in terms of their algorithmic properties. 2) Showing how the derivation of ever more powerful RL algorithms displays a trend towards BBO. 3) Continuing this trend by applying two modifications to the state-of-the-art PI 2 algorithm, which yields an algorithm we denote PI BB. We show that PI BB is a BBO algorithm. 4) Demonstrating that PI BB achieves similar or better performance than PI 2 on several evaluation tasks. 5) Analyzing why BBO outperforms RL on these tasks. Rather than making the case for BBO or RL – in general we expect their relative performance to depend on the task considered – we rather provide two algorithms in which such cases can be made, as the algorithms are identical in all respects except in being RL or BBO approaches to policy improvement. 1

Read the paper · More papers on PaperTik