AdaMoT: Adaptive Motion-Aware Transformer for Efficient Visual Tracking

Yongjun Wang, Xiaohui Hao · IEEE Signal Processing Letters · 2025

Visual object tracking utilizing adaptive computation presents challenges stemming from the complexities of modeling intricate motion patterns and achieving computational efficiency. While recent transformer-based trackers have shown promising results, they struggle to effectively capture varying motion dynamics and often waste computation on less informative regions, leading to degraded performance under fast motion and occlusion. In this letter, we present AdaMoT, an innovative motion-aware transformer framework featuring three lightweight modules that integrate adaptive attention and motion estimation: a Lightweight Adaptive Motion Estimation (LAME) module that guides transformer attention through motion pattern modeling, a Saliency-based Hard Attention Sampling (SHAS) module that reduces computation by 60% through focusing on motion-critical regions, and an Adaptive ViT Attention Head Adjustment (AVAHA) module that dynamically allocates attention heads based on motion complexity. Our framework uniquely integrates motion estimation with transformer attention through a shared feature space, achieving robust tracking with minimal overhead. Comprehensive testing indicate that AdaMoT attains superior performance on various demanding benchmarks (75.1% AO on GOT-10 k, 84.9% AUC on TrackingNet, 72.9% AUC on LaSOT) while maintaining real-time speed (32.1 FPS) with only 4% FLOPs increase.

Read the paper · More papers on PaperTik