Multi-Level Features Fusion for Zero-Shot Object Pose Estimation
Nanxin Huang, Chi Xu · 2025
Driven by advancements in industrial production and artificial intelligence, the need for pose estimation of new ob-jects in areas like robotic manipulation and virtual reality is increasing. We introduce a zero-shot object pose estimation approach that identifies the poses of objects excluded from the training dataset, removing the requirement for re-modeling. The method is built around a multi-level features fusion framework de-signed to enhance generalization. First, a trainable feature extraction module filters and selects multi-level features extracted by the backbone network. Unlike traditional convolutional ker-nels, we incorporate a dynamic convolution kernel to enhance the feature extraction capability. Second, in the feature fusion module, we adopt a dynamic weight generation strategy to perform weighted fusion of multi-level features. This method enhances template matching by effectively describing similarities between unseen objects (those absent from the training set) and templates, leveraging robust and adaptive feature representations to narrow the gap with seen objects. Experimental results demonstrate that our approach achieves state-of-the-art performance on two popu-lar benchmark datasets, LineMod and LineMod-Occlusion, proves that our method has better generalization than previous models.