ACT: an ACTNet for visual tracking
Ning Li, Qingge Ji, Tianjun Ma · IET Image Processing · 2018
Owing to convolutional neural network (CNN) models’ success in various fields of computer vision, the authors proposed an advanced convolutional network (ACTNet) to enhance the accuracy of visual tracking. Different from prior methods, they regard a CNN as not only a semantic feature map extractor but also a position predictor. Rectified Linear Unit (RLU) and sigmoid are both used in ACTNet for feature extraction and position determination. To avoid overfitting in pre‐training, they introduce adding Erlang noise to create more training samples and to improve the robustness of each base learner. Experiments on widely used evaluation datasets demonstrate that their proposed ACT method outperforms state‐of‐the‐art methods.