Improving Vision-Language Models With Attention Mechanisms for Aerial Video Classification
Nguyen Anh Tu, Nartay Aikyn · IEEE Geoscience and Remote Sensing Letters · 2025
Vision-language models (VLMs), particularly contrastive language-image pretraining (CLIP), have recently demonstrated great success across various vision tasks. However, their potential in aerial video understanding, an increasingly active area of remote sensing (RS), remains underexplored. This is due to challenges posed by aerial data, such as UAV movement, extreme camera angles, and complex spatiotemporal dependencies. To tackle these challenges, we propose an effective method called CLIP-AVC, which adapts CLIP to classify aerial videos into predefined classes. Specifically, we leverage CLIP’s multimodal transferability by utilizing its encoders to extract robust visual and textual features. We then employ a temporal transformer to capture the interactions among the visual features. To address the lack of inductive bias in the CLIP’s visual encoder, we integrate the temporal transformer’s outputs with 3-D features using a cross-transformer, thereby allowing the spatiotemporal locality of aerial videos. In addition, existing methods often fail to explore the semantic alignment between classes and video features. To further overcome these limitations, we propose a context-enriched transformer that employs self-attention mechanisms to adaptively refine visual and textual representations. Experimental results on two benchmark datasets validate the robustness of CLIP-AVC, demonstrating its potential to significantly advance VLMs for aerial scene understanding.