Semantic-Guided Mamba Fusion for Robust Object Detection of Tibetan Plateau Wildlife
Ping Lan, Yukai Xian, Te Shen, Yurui Lee, Qijun Zhao · Electronics · 2025
Accurate detection of wildlife on the Tibetan Plateau is particularly challenging due to complex natural environments, significant scale variations, and the limited availability of annotated data. To address these issues, we propose a semantic-guided multimodal feature fusion framework that incorporates visual semantics, structural hierarchies, and contextual priors. Our model integrates CLIP and DINO tokenizers to extract both high-level semantic features and fine-grained structural representations, while a Spatial Pyramid Convolution (SPC) Adapter is employed to capture explicit multi-scale spatial cues. In addition, we introduce two state-space modules based on the Mamba architecture: the Focus Mamba Block (FMB), which strengthens the alignment between semantic and structural features, and the Bridge Mamba Block (BMB), which enables effective fusion across different scales. Furthermore, a text-guided semantic branch leverages knowledge from large language models to provide contextual information about species and environmental conditions, enhancing the consistency and robustness of detection. Experiments conducted on the Tibetan wildlife dataset demonstrate that our framework outperforms existing baseline methods, achieving 70.2% AP, 88.7% AP50, and 76.8% AP75. Notably, it achieves significant improvements in detecting small objects and fine-grained species. These results highlight the effectiveness of the proposed semantic-guided Mamba fusion approach in tackling the unique challenges of wildlife detection in the complex conditions of the Tibetan Plateau.