Depth-Aware Object Tracking With a Conditional Variational Autoencoder

Wenhui Huang, Jason J. Gu, Yinchen Guo · IEEE Access · 2021

Object tracking is a fundamental task in computer vision and artificial intelligence. However, state-of-the-art object tracking approaches are still prone to failures and are imprecise when applied to challenging scenarios, and their results are generally confidence agnostic. An imprecise deterministic output with low confidence may lead to disastrous consequences and a lack of proof for subsequent operations and human interventions. Deep network training with ambiguous data or the noise inherent in observations (i.e., data uncertainty or aleatoric uncertainty) will result in inherent uncertainties in predictions. In this paper, we exploit probabilistic depth-aware object tracking with a conditional variational autoencoder (CVAE). First, we build a bridge between the Siamese network and the variational autoencoder conditioned with depth images and propose a novel multimodal Bayesian object tracking method. Second, our proposed method yields a complete probability distribution that enables the production of multiple plausible features. Third, the variational autoencoder conditioned by depth images encodes a low-dimensional latent space that conducts depth-aware tracking, which has obvious advantages for challenging tracking scenarios. Our proposed tracking method outperformed the state-of-the-art trackers on the VOT 2016, VOT 2018, and VOT 2019 datasets.

Read the paper · More papers on PaperTik