Structured reward functions using STL

Anand Balakrishnan, Jyotirmoy V. Deshmukh · 2019

In this work we present a new method for shaping reward functions to train reinforcement learning agents using signal temporal logic (STL) formulas. The proposed approach uses the robustness metric of partial signal traces against STL specifications to generate locally shaped rewards, doing this in a manner that is agnostic of the learning algorithm used by the reinforcement learning agent.

Read the paper · More papers on PaperTik