Multiparametric Analysis of Multi-Task Markov Decision Processes: Structure, Invariance, and Reducibility

Jae-Uk Shin, Insoon Yang · IEEE Control Systems Letters · 2024

Modern sequential decision-making problems such as robotic control require on-the-fly adaptation to a reward function encoding a novel task, particularly when there is no luxury of performing costly online planning procedures. In this article, we present a multiparametric linear programming (mp-LP) approach that addresses this problem. The key idea is to characterize the optimal return of a multi-task Markov decision process as the optimal value of an mp-LP problem having reward functions (or vectors) as parameters. This mp-LP characterization enables the simultaneous computation of the optimal returns for multiple reward functions. Our multiparametric analysis provides geometric structures of the mp-LP problem by essentially identifying all polyhedral sets of reward functions that lead to the same optimal policy. This allows us to identify the types of reward function transformations that guarantee the invariance of optimality, effortlessly yielding the potential function-based reward shaping technique as a corollary. Finally, our analysis enables us to address the reducibility problem that seeks the minimal dimensions to capture essential information about any reward function.

Read the paper · More papers on PaperTik