Foundation Model Vision Transformers are Great Tracking Backbones

Tristan Kenneweg, Philp Kenneweg, Barbara Hammer · 2024

The recent breakthroughs in foundation models for image processing [1], [2] have made using Vision Transformer embeddings for downstream image tasks a great option in many applications. However, best practices on how to use these embeddings for a given task have not been established yet.In this paper we investigate the suitability of foundation models for the single object tracking task. We do this by developing and implementing a zero-shot patch tracking method and a deep learning system which builds upon foundation model Vision Transformer embeddings. We evaluate these methods on the challenging GOT-10k dataset and show favorable results compared to those obtained with an established ResNet-50 backbone. We make our code publicly available as a GitHub repository, alongside videos that show our tracking results.

Read the paper · More papers on PaperTik