End-to-End Binocular Active Visual Tracking Based on Reinforcement Learning

Biao Zhang, Songchang Jin, Shaowu Yang, Qianying Ouyang, Yuxi Zheng, Dianxi Shi · 2023

Active Visual Tracking (AVT) aims at controlling the motion system of the tracker to follow the target given visual observations. Existing works have achieved significant breakthroughs by leveraging deep reinforcement learning. However, these methods rely solely on a single visual observation image, and the tracker can only infer relationship between the position of itself and the target according to the size of target in observed images, lacking ability to perceive the distance of the target. Furthermore, environmental background interference poses challenges for the tracker. To addresses these issues, we propose a binocular stereo matching method for AVT. The disparity feature of the target is obtained by stereo matching of the binocular images to perceive the distance of target. To address the problem of background interference, we introduce template matching into active visual tracking to embed target template information into the state representation, thereby enhancing the ability of the tracker to identify objects. To ensure stable policy training, we adopt an asymmetric training mechanism that utilizes true state information during training, further improving tracking performance. Experimental results show that our method achieves better results in different test scenarios compared to baselines.

Read the paper · More papers on PaperTik