Transformer Tracker by Attention Feature Fusion Module and Online Update Module
Xiaohan Liu, Aimin Li, Deqi Liu, Dexu Yao, Mengfan Cheng · 2022
In recent years, Transformer has shown strong competitiveness and better performance in the field of object tracking. However, during tracking, object occlusion and similar objects can affect the tracking effect. If only features of the last layer are used, sufficient detail information cannot be obtained, and it is difficult to accurately locate the object. As the ResNet-50 network gets deeper and deeper, feature maps of the last layer obtain high-level semantic information, but the feature maps have low resolution, poor ability to perceive details, and rough locations. The low-level feature maps have high resolution and contain more detailed feature and location information, but the semantic information is not strong. Second, if only the template features of objects in the first frame are learned, factors such as occlusion, deformation, and complex background can easily lead to tracking failure. To address these issues, we propose a Transformer tracker by attention feature fusion module and online update module. Firstly, an attention feature fusion module is designed, which can fully integrate low, medium and high-level features, obtain rich detailed information and semantic information, and improve the ability of features to express objects. Subsequently, we use the transformer structure and add an online update branch to solve the tracking failure caused by only learning the template of first frame and the drift caused by the accumulation of updates. Experimental results show that our algorithm outperforms several well-known tracking algorithms in terms of tracking accuracy and robustness.