A Multimodal Spatial Non-Cooperative Target Intent Recognition Method Based on Dynamic Attention and Cross-Modal Interaction
Tian Ma, Ren Wanzhu, Shang Chuyang, Ma Shaofeng, Jiahui Li · IFAC-PapersOnLine · 2025
With the increasing number of space activities, intelligent perception and intent recognition of non-cooperative space targets are of great significance to space security. However, existing methods primarily face two key challenges: on the one hand, unimodal perception struggles to effectively capture the fine-grained positional changes and dynamic behavioral characteristics of space targets; on the other hand, traditional algorithms inadequately utilize multimodal data collaboratively, resulting in limited predictive capability for target intent. To address these issues, this paper proposes a multimodal space non-cooperative target intent recognition method named FDC-CLIP based on dynamic attention and cross-modal interaction. First, we design an intent-driven dynamic visual attention module (DVAM). By encoding high-level semantic instructions into low-dimensional query vectors through an intent-aware feature extractor, it achieves task-adaptive selection of key regions. Second, we propose a cross-level residual interaction mechanism. Based on the "CLS-Patch-Intent" triangular feature interaction framework, this mechanism simultaneously preserves global semantic information and local dynamic features through differentiable residual connections. Finally, we introduce a multi-head cross-attention fusion (EMCAF) module. This module employs multi-head attention to capture diverse interaction patterns between image and text features, while cross-attention further enhances the alignment capability between visual and textual features, thereby enabling more accurate modeling of fine-grained behavioral characteristics of targets. The experimental results show that on the self-built multimodal dataset, the cross-modal feature alignment of the proposed method outperforms the existing baseline model, and the recognition accuracy is improved by 3.89%, 10.56% for CLIP, and 23.54% for MobileCLIP compared to ResNet50, and it is able to better adapt to the requirements of intent recognition for spatial non-cooperative targets.