Robot Skill Learning: From Reinforcement Learning to Evolution Strategies

Freek Stulp, Olivier Sigaud · Paladyn Journal of Behavioral Robotics · 2013

Abstract Policy improvement methods seek to optimize the parameters of a policy with respect to a utility function. Owing to current trends involving searching in parameter space (rather than action space) and using reward-weighted averaging (rather than gradient estimation), reinforcement learning algorithms for policy improvement, e.g. PoWER and PI

Read the paper · More papers on PaperTik