MV-Poseformer: Direct Multi-View 6D Object Pose Estimation with Transformers
Chao Wen, Yaling Liang · 2025
In this paper, we focus on the task of 6D object pose estimation from multi-view RGB images, and introduce a novel framework, termed as MV-PoseFormer, which leverages transformers to aggregate multi-view information and directly regress object poses, without any post-optimization. More specifically, MV-PoseFormer extracts features from multi-view RGB images and realizes direct pose estimation using two carefully designed modules, i.e., Multi-view Pose Estimator and Residual Pose Regressor, based on visual transformers. Multi-view Pose Estimator models camera poses of images as augmented encodings to inject spatial context, and equips pose tokens to be jointly optimized from multi-view features via transformers for regressing pose estimates. Residual Pose Regressor further refines the object poses by using the coarse ones to sample features of keypoints from multi-view features and employing transformers to strengthen pose tokens, from which the residual poses are regressed for more precise predictions. We conduct experiments on the StereOBJ-1M and YCB-V datasets to evaluate the effectiveness of our proposed MV-PoseFormer, which outperforms the existing methods, including those with timeconsuming multi-view post-optimization. Comprehensive ablation studies are also conducted to verify the advantages of individual designs in MV-PoseFormer.