Generating Textual Descriptions for Remote Sensing Images
Maryam Mehmood, Ahsan Shahzad, Amin Ullah, Farhan Hussain, Samra Nawazish, Ruhul Amin Khalil · IEEE Access · 2026
Generating descriptive and semantically meaningful captions for remote sensing images remains a challenging task due to the inherent complexity of aerial scenes, characterized by diverse spatial arrangements, scale variations, and ambiguous semantics. To address these challenges, this study introduces a hybrid deep learning framework that integrates a Convolutional Neural Network (CNN), a Transformer encoder, and a Long Short-Term Memory (LSTM) decoder. In the proposed architecture, ResNet50 is employed to extract hierarchical visual features, while the Transformer encoder models long-range spatial dependencies and global contextual relationships. The resulting feature representations are refined through a fusion mechanism and subsequently utilized by an LSTM-based decoder to generate coherent and context-aware textual descriptions. The effectiveness of the proposed approach is evaluated on benchmark remote sensing datasets, namely UCM and RSICD. Experimental results indicate that the model achieves competitive performance, with BLEU-4 scores of 0.74 and 0.54, and CIDEr scores of 3.32 and 3.49, respectively. These findings demonstrate the capability of the proposed framework to generate accurate and contextually relevant captions for complex aerial imagery. The study highlights the potential of hybrid architectures in improving automated remote sensing image understanding and caption generation.