Integrated Framework for Remote Sensing Image Captioning Using RSIFE_Net and Transformer Model

International journal of intelligent engineering and systems · 2025

The growing availability of remote sensing imagery has heightened the need for automated methods to extract meaningful awareness from large volumes of data.Image captioning, which merges computer vision and natural language processing, offers a promising solution to interpret and describe remote-sensing images in a humanlike manner.This research introduces an integrated framework that combines the Remote Sensing Image Feature Extraction Network (RSIFE_Net) with transformer models to increase the accuracy and efficiency of remote sensing image captioning.The RSIFE_Net extracts discriminative visual features.These features are then processed by a transformer-based captioning model to generate logical and meaningful captions.The transformer architecture effectively captures long-range dependencies and semantic relationships.Four datasets were employed to evaluate the proposed image captioning framework: RSICD, UCM-Captions, Sydney-Captions, and NWPU-Captions.The model achieved a BLEU-4 score of 0.5525 and a CIDEr score of 3.0404 on the RSICD dataset, significantly outperforming previous models.It also performed better on the UCM-Captions dataset, with a BLEU-4 score of 0.7722.On the NWPU-Captions dataset, all scores are better than the classical models including BLEU-4: 0.6782 and a CIDEr score of 2.3598.These results confirm that the proposed method effectively enhances remote sensing image captioning tasks.While qualitative measurements indicate the model's ability to capture important visual features and spatial relationships, quantitative measurements demonstrate the model's competitive performance compared to state-of-theart methods.

Read the paper · More papers on PaperTik