HyPoseNet: a hybrid CNN for 6D pose estimation with point feature convolution and attention fusion
Allison Jin, Hai Wang, Chunlai Yang, Junhao Wen, Otoide Iretiose Belief, Luyando Stanley Kabalata, Jinsong Gui · Engineering Research Express · 2025
Abstract The application of RGB-D image data in tasks such as intelligent perception and pose estimation for robots has recently garnered significant attention. A major challenge is effectively utilizing the complementary modalities in RGB-D data sources. A 6D pose estimation method based on the multimodal fusion method, with a hybrid convolutional neural network (CNN) architecture integrating point feature convolution and attention mechanism is proposed in this study. The method integrated a two-stage framework for segmentation and pose regression of RGB-D data, effectively extracting target features and enhancing the model’s interpretability and robustness. A convolutional block attention module is introduced to recalibrate the feature maps dynamically, which focuses on information-rich regions and channels, thereby improving the overall performance. A series of comprehensive evaluations were carried out on standard benchmark datasets to validate the effectiveness of the proposed method. On the Occlusion LineMOD dataset, it achieved an average ADD accuracy of 96.9%. For the YCB-Video dataset, the method reached a 94.6% average accuracy under the ADD-S AUC metric, and 97.7% of predictions yielded ADD-S distances below 2 cm. These experimental outcomes clearly illustrate the robustness and high precision of the proposed approach.