RSIC-GMamba: A State-Space Model With Genetic Operations for Remote Sensing Image Captioning

Lingwu Meng, Jing Wang, Yan Huang, Liang Xiao · IEEE Transactions on Geoscience and Remote Sensing · 2025

Recent advancements indicate that the novel Mamba framework, with its linear reasoning capabilities and performance on par with the Transformer framework, has become a popular alternative to Transformers. However, there is still a lack of exploration in remote sensing image captioning (RSIC). One important challenge is that the 1-D selective state-space model (SSM) is not suitable for 2-D remote sensing images (RSIs). To mitigate this challenge, this article proposes an SSM with genetic operations for RSIC (RSIC-GMamba) that integrates the search capabilities of heuristic genetic algorithm and the global modeling capabilities of Transformer into the Mamba framework. To comprehensively capture the multiscale global context of RSIs, we design a genetic SSM by incorporating a dilated convolution, genetic operations (crossover and mutation), and self-attention mechanism into the selective SSM. Specifically, dilated convolutions are employed to extract multiscale visual features through varying dilation rates. Crossover operation expands the spatial arrangement of image regions to thoroughly capture contextual information and mutation operation introduces randomness to enhance the model’s robustness. The self-attention mechanism is integrated to model relationships among SSM hidden states, thereby enhancing visual context. In addition, to fully utilize high- and low-level visual-semantic information, we propose a vision and scene text aggregation (ViSTA) module based on gating mechanisms. Experimental results on four RSIC datasets demonstrate the effectiveness of the proposed RSIC-GMamba. The code will be publicly available athttps://github.com/One-paper-luck/RSIC-GMamba.

Read the paper · More papers on PaperTik