DTF-Net: A Dual Attention Transformer-based Fusion Network for 6D Object Pose Estimation
Tao An, Kun Dai, Ruifeng Li · 2024
In the field of computer vision, 6D object pose estimation has emerged as a crucial research direction. Recent advancements leverage deep learning networks that integrate RGB images with depth images for enhanced feature learning. In this study, we introduce DTF-Net, a dual attention transformer-based fusion network for 6D object pose estimation using RGBD data. By combining appearance features from RGB data with geometric features from depth images, our method comprehensively captures object characteristics in a scene, significantly enhancing pose estimation accuracy. Central to our approach is the dual attention transformer (DAT) based fusion module, which fuses these two complementary features. The DAT surpasses traditional transformer by accurately capturing dual-modal data features and reducing computational complexity. Our network’s superior performance is validated through experiments on the YCB-Video dataset, where it outperforms current state-of-the-art models.