Semantic Object Navigation With Segmenting Decision Transformer
Aleksei Staroverov, Tatiana Zemskova, Dmitry A. Yudin, Aleksandr I. Panov · IEEE Access · 2025
Object navigation remains a fundamental challenge in robotics, particularly when agents must reach targets specified by semantic categories. While existing approaches often treat semantic understanding and navigation as separate components, we demonstrate that their tight coupling is crucial for robust performance. We present SegDT (Segmenting Decision Transformer), a novel architecture that jointly learns to predict semantic segmentation masks and navigation actions through a unified transformer-based model. Our key insight is that temporal information from sequential observations can simultaneously enhance both segmentation quality and navigation decisions. To address the inherent challenges of transformer-based navigation—notably poor sample efficiency and computational complexity—we introduce a two-phase training approach: offline pretraining on expert demonstrations followed by online policy refinement through knowledge transfer from a recurrent neural network. Extensive experiments in the Habitat simulator demonstrate that SegDT achieves higher results using predicted segmentation masks, outperforming a single-frame baseline with a pre-trained semantic segmentation model and approaching the performance of systems using ground truth semantic information. Our ablation studies reveal that SegDT’s temporal processing also improves segmentation quality, highlighting the synergistic benefits of joint optimization. When integrated into complete object navigation systems, SegDT enhances overall performance by 9.6% in path efficiency compared to the state-of-the-art method. The code of SegDT is made publicly available athttps://github.com/CognitiveAISystems/SegDT