AquaSOT: end-to-end dual-modal single object tracking in underwater environments
Luhang Hong · DR-NTU (Nanyang Technological University) · 2026
Underwater environments pose significant challenges for visual object tracking. Turbidity, light absorption, and scattering severely degrade RGB cameras, limiting their effective range to mere meters in adverse conditions. Forward-Looking Sonar (FLS) complements these limitations through acoustic imaging that remains robust regardless of optical conditions, though at lower spatial resolution. This complementarity motivates dual-modal tracking, yet combining these fundamentally different sensing modalities is difficult: the sensors produce distinct visual representations with different spatial coverage, and direct feature alignment between optical and acoustic imagery is impractical. This dissertation addresses dual-modal single object tracking in underwater environments, where RGB and sonar streams provide complementary but spatially misaligned observations of the same target. The proposed framework, AquaSOT, employs a dual-stream Vision Transformer (ViT) architecture with two separate ViT-Base backbones for modality-specific feature extraction. A bidirectional Cross-Modal Attention (CMA) module enables each modality to draw on the other's strengths: RGB features transmit fine-grained texture and boundary information that improves sonar's coarse spatial predictions, while sonar features provide stable acoustic localization cues that anchor RGB tracking when underwater optical conditions deteriorate. This cross-modal correspondence is learned end-to-end without requiring explicit sensor calibration. An asymmetric fusion weight strategy applies full fusion during training to thoroughly learn cross-modal projections, then reduces fusion strength at inference to prevent degraded features from one modality from corrupting the other, thereby preserving each sensor's distinctive tracking capability. Separate prediction heads accommodate the different spatial coverage of the two sensors, while confidence-aware template updating and dual-modal target loss detection maintain robustness when either modality temporarily fails. Experimental evaluation on the RGB-Sonar tracking benchmark demonstrates that AquaSOT achieves an overall Success AUC of 0.7276 and Precision@20 of 0.9721. Per-modality analysis confirms genuine complementarity: RGB provides precise boundary estimation (Success AUC 0.8429) while sonar delivers stable center localization (Precision@20 1.0000, center error 4.77 pixels) with continuous target visibility. Ablation studies validate the contribution of each component: dual-modal fusion improves over RGB-only tracking by 5.9% and over sonar-only tracking by 29.9% in Success AUC, while disabling CMA fusion alone degrades performance by 16.8%, confirming cross-modal attention as the primary driver of improvement. These results demonstrate that learned cross-modal attention between optical and acoustic modalities significantly enhances underwater tracking robustness, establishing a practical foundation for autonomous underwater vehicle navigation and marine research applications.