Remote Sensing Image Captioning using CNN and LSTM
Vaishnavi T V, C Indu · 2024
Image captioning is a rapidly expanding field in computer vision research and it involves generating a comprehensive description for an input aerial image. The community has been paying more attention to it lately because it offers additional semantic information about images. To properly convey the connections between the objects and the features included in remote sensing pictures with their captions, this work proposes a system for captioning those images. The system is designed to generate and exploit textual descriptions. For captioning images, an encoder-decoder structure is employed. Convolutional neural networks (CNNs) are first used to encode the image’s visual properties. After that, the encoded features are sent to a language model, or decoder. For word by word descriptions of the image content, recurrent neural networks (RNNs), often referred to as long short term memory (LSTM), are commonly employed as language models. We combined the Bahdanau attention model with LSTM to enable learning to be focused on a specific region of the image to enhance performance. Four different pre-trained CNNs were evaluated for model performance in this study. Experimental results from UAVIC dataset are presented.