Robust Visual Tracking Based on Adversarial Fusion Networks
Ximing Zhang, Mingang Wang, Jinkang Wei · 2018
Recent advances in visual tracking showed that deep Convolutional Neural Networks(CNN) trained for image classification can be strong feature extractors for discriminative trackers. However, due to the long-tail of categories, occlusions, deformations and some other attributes are so rare that they will hardly happen. Yet we want to learn a model invariant to such occurrences for fear that we could handle complex attributes in tracking procedure. In this paper, we propose an alternative solution. We propose to learn an Adversarial Fusion Networks(AFN) that generates examples with occlusions and deformation based on the internal structure of Region Proposal Network (RPN). We discovered that the internal structure of Adversarial Fusion Networks(AFN)'s top layer feature can be utilized for robust visual tracking. We illustrated that such networks can be more robust when the tracking object suffering from occlusion and deformation. Without ensemble and any extra treatment on feature maps, our proposed method achieved state-of-the-art results on several large scale benchmarks including OTB50, OTB100 and VOT2016.