Leveraging privileged information and symmetry for policy learning in partially observable robotic systems
Hải Thanh Nguyễn · 2024
Partially observable systems are common in robotics, yet much of the existing research concentrates on fully observable environments, where agents unrealistically have complete access to state information. The key challenge in policy learning under partial observability lies in learning policies that not only optimize rewards but also incorporate information-gathering actions to acquire more knowledge about the environment, all while retaining and recalling critical information for future decision-making. This presents a far more complex problem compared to fully observable settings, where the agent's task is solely to optimize for rewards without the added burden of handling uncertainty and incomplete information. This dissertation addresses these challenges by assuming the availability of privileged information during training, such as the state distribution given historical observations and actions, fully observable solutions, or knowledge of domain symmetry. We propose novel methods that leverage these forms of privileged information to accelerate policy learning in partially observable environments. Our methods enhance the agent's ability to gather and utilize information while improving its capacity to retain and recall key details crucial for decision-making. These approaches significantly outperform existing baseline methods. Additionally, the policies learned under these settings can be transferred to work well with real-world hardware without additional training, demonstrating their practical applicability. The proposed methods pave the way for building more intelligent robotic systems capable of reasoning and acting under uncertainty in realistic environments where full observability is not possible.--Author's abstract