FA-transformer: a remote sensing image captioning framework integrating multiscale features and enhanced attention mechanisms
Weiyi Wang, Guoqing Zhou · 2025
Remote Sensing Image Captioning (RSIC) is designed to automatically convert satellite or aerial imagery into coherent textual descriptions, allowing non‑specialist users to rapidly grasp complex visual content. Nevertheless, the wide range of object scales and the abundance of fine‑grained local features in such imagery present serious obstacles for traditional CNN and RNN–based techniques. To address these shortcomings, we introduce the FA‑Transformer framework: it employs an Atrous Spatial Pyramid Pooling (ASPP) module to capture multi‑scale features, an Enhanced Attention Mechanism (EAM) that fuses CBAM with CoordAtt for adaptive channel and spatial weighting, and a Transformer encoder augmented with CLIP image embeddings to strengthen semantic understanding. Experimental evaluations show that FA‑Transformer considerably surpasses current state‑of‑the‑art methods in producing precise and informative captions for remote sensing imagery.