SamPose: Generalizable Model-Free 6D Object Pose Estimation via Single-View Prompt
Wubin Shi, Shaoyan Gai, Feipeng Da, Zeyu Cai · IEEE Robotics and Automation Letters · 2025
Object pose estimation in open-world scenarios is a critical challenge in robotics, virtual reality, and autonomous driving. In this letter, we introduce SamPose, a novel framework designed to achieve model-free 6DoF pose estimation of any target object in open-world environments using only a single-view prompt. SamPose consists mainly of an Open-world Object Detector (OOD) and a Coarse-to-Fine Pose Estimator (CFPE). The OOD utilizes a pre-trained EfficientSAM model to perform zero-shot segmentation matching tasks. It selects the proposals most similar to new objects based on matching scores derived from semantic, geometric, and local descriptors. In the CFPE phase, a sparse keypoint matcher, guided by DINO semantics, first performs robust keypoint matching and calculates an initial pose. Then, after aligning the perspectives from two views, a two-stage semi-dense keypoint matcher is used to compute reliable point correspondences and ultimately determine the object's pose. Finally, our extensive experiments demonstrate its robustness and competitive performance.