Efficient Feature Fusion Network with Attention for Visual Object Tracking

Yanxia Wei, Deying Feng, Jian Mu, Chongjing Wang · 2025

Recently, Siamese-based tracking methods have emerged as well-liked approaches in the field of visual object tracking. Most existing Siamese-based tracking methods usually employ the cross-correlation operation to learn the principal information of target template and search regions. However, because of the excessive emphasis on calculating the similarity scores between them, the cross-correlation strategies often ignore either the local information or the semantic information of the target. To handle this issue properly, we present an efficient feature fusion network with attention for visual object tracking. Firstly, we utilize the Channel-wise Augment and the Pixel-wise Augment to extract the semantic features and the local spatial features from the template branch and the search branch simultaneously. Then, we employ the Attention Augment to fully fuse the two kinds of feature information to augment the power of feature representation capability. We rigorously evaluate our tracker through the comprehensive experiments on several widely-used benchmarks, ensuring fair comparisons with existing methods. Extensive experiments are conducted to demonstrate that our proposed method surpasses the superior performance in comparison with the current state-of-the-art tracking algorithms.

Read the paper · More papers on PaperTik