Remote Sensing Image Captioning with SVM-Based Decoding
Genc Hoxha, Farid Melgani · 2020
With the fast development of remote sensing (RS) technology we are now able to acquire high resolution images. To cope with the new challenges of analyzing such images, a recently introduced tool is RS image captioning (IC). With respect to conventional techniques such as scene classification, RSIC provides more information about an image. It aims to generate a description that summarizes the content of an image. Most of RSIC systems are based on deep learning frameworks (encoder-decoder). The performance of these frameworks strongly depend on the number of annotated samples used during training. In this paper, we propose an alternative RSIC system for a relatively small dataset based on support vector machines (SVMs). A pre-trained CNN is used to extract the image visual features and a network of SVMs is used to generate the descriptions. Experimental results on a RS image archive composed of images acquired by unmanned aerial vehicles (UAV), show that the proposed IC system could be an interesting alternative to deep learning frameworks when only small training samples are available.