Augmented Hierarchical Scene Prior Learning With Context-Based Scene Completion Network for Visual Semantic Navigation

Jiaxu Kang, Chengyang Zhu, Bolei Chen, Ping Zhong, Haonan Yang, Tao Zou · IEEE Transactions on Cognitive and Developmental Systems · 2025

The Visual Semantic Navigation(VSN) requires the agent to navigate to a target object of specified category in a previously unseen scene. To tackle this task, the agent must learn a nimble navigation policy by utilizing spatial patterns and semantic co-occurrence relations among objects in the scene. Prevailing approaches extract scene priors from the instant visual observations and solidify them in neural episodic memory or explicit scene representations to achieve flexible navigation. However, due to the oblivion and underuse of the scene priors, these methods are plagued by repeated exploration, effective knowledge sparsity, and wrong decisions. To alleviate these issues, we propose a novel VSN policy, HSPNav, based on Hierarchical Scene Priors (HSP) and Deep Reinforcement Learning (DRL). The HSP contains three components, i.e., the egocentric semantic map-based Local Scene Priors (LSP), the commonsense relational graph-based Global Scene Priors (GSP), and the Attentionbased retrieval mechanism that retrieves conducive contextual memories closely related to the immediate LSP from the GSP. Furthermore, we propose a Semantic Map Completion Network with Context Association Exploitation module(CAE-SMCN) to exploit the context associations for unobserved scene inference on the egocentric map, resulting in augmented LSP. Extensive experiments on MP3D and HM3D show that our HSP facilitates the scene priors and navigation policy learning, and outperforms the existing methods. Finally, we implement the sim-to-real transfer for the navigation policy and demonstrate that HSPNav can generalize well to realistic settings.

Read the paper · More papers on PaperTik