An End-to-End Location and Regression Tracker with Attention-based Fused Features
Qinyi Zhang, Shishuai Du, Huihua Yang · 2019
Visual object tracking, which continuously generating bounding box of a specific target initialized in the first frame in a video, is one of the most fundamental and challenging tasks in computer vision. In this paper, we derive a tracker that combines correlation filter and siamese network, which both achieve superior performance and complement each other, jointly improving the tracker's performance. However, this tracker still has some problems common to most state-of-the-art trackers, that is, discrimination of features is not strong enough, and the scale estimation method is limited. To solve these problems, we propose LR-AFNet, an end-to-end tracker for target location (via the siamese network) and bounding box regression, to better locate and estimate scales, and using attention-based multi features fusion method to extract more discriminative features. Firstly, the feature extraction backbone network is redesigned to improve the location accuracy. Secondly, in order to obtain more effective features, deep and shallow features are fused with an attention mechanism, which adaptively learn the fusion weights. Finally, the bounding box regression branch is added to the siamese similarity network, forming an end-to-end trained framework. Experiments proved that our proposed method achieves competitive performance on several benchmark datasets, compared with the state-of-the-art.