Reinforcement learning in partially observable environments using approximate information state

Amit Kumar Sinha · 2021

In this thesis, reinforcement learning (RL) for multi-stage decision making problems where the decision maker or the agent has partial observations of the state are considered. Many real world systems are partially observable---some examples are autonomous driving, robotics, smart grids, etc. However, most of the research on RL is centered around fully observable systems. Partially observable systems are considerably harder than fully observable systems because finding the optimal policy involves considering the entire trajectory of past actions and observations to decide on the current action. The exponential growth of history-dependent policies makes it much harder to find an optimal policy, whereas in fully observable systems the policy search space is much smaller. This further exacerbates issues related to high dimensionality of the observation, action and/or state spaces. The belief distribution over states can be used instead of the past actions and observations to circumvent the exponential growth of policies, but this requires model information about the system which may not always be available in a reinforcement learning setting.An alternative to the history or belief distribution over states is the more general notion of the information state which is a representation that is sufficient for learning optimal agent behavior. The entire history and the belief state may be viewed as instances of information state. A desirable information state allows us to compress the history to a smaller size so that the policy search space is smaller. An approximate information state (AIS) can be learnt from data using function approximation without requiring any model information.One of the most important steps in learning an effective AIS is being able to predict the probability distribution of the next AIS/observation given the current AIS and action. This is done by optimizing integral probability metrics (IPMs) which try to decrease the ``distance'' between the predicted distribution and the actual distribution. We demonstrate the versatility of AIS using two different choices of IPMs. One is the Wasserstein distance which is optimized in terms of a surrogate KL-divergence loss (KL IPM). The second is a distance-based maximum mean discrepancy loss (MMD IPM).We show the scalability of the proposed methods through numerical experiments in environments of increasing difficulty. These methods are compared with approximate planning solutions for low and moderate-dimensional environments for a model-based baseline. For a model-free RL baseline, a recently proposed method using proximal policy optimization (PPO) with recurrent connections (LSTM) is used for all environments. We show that the proposed RL methods based on AIS outperform the PPO with LSTM baseline in most environments. We also show that the performance achieved by our RL algorithms is close to the near optimal approximate planning solutions in the low and moderate-dimensional environments

Read the paper · More papers on PaperTik