Safe Policy Iteration – Supplementary Material

Matteo Pirotta, Marcello Restelli, Alessio Pecorino, Daniele Calandriello · 2014

This document provides additional material to the main paper. In particular, it provides: 1) the complete set of theorems, lemmas and corollaries with the relative proofs; 2) additional experiment in chain walk and BlackJack domains; 3) a detailed analysis of the performances in terms of computational time. 1. Proofs In this section, we will prove the lemmas, theorems, and corollaries stated in our paper. Lemma 3.1 Let π and π ′ be two stationary policies for an infinite horizon MDP M with state transition matrix P. The L1–norm of the difference betweentheir γ–discountedfuture state distributionsunder startingstate distribution µ can be upper bounded as follows: ∥d π′ µ −d π µ

Read the paper · More papers on PaperTik