Multimodal Learning for Earth Observation: Automating Satellite Image Captioning with Geo-FMs
Chiarabini, Luca, Espinoza Molina, Daniela, Zappacosta, Antony, Kuzu, Ridvan Salih, Camero, Andres · elib (German Aerospace Center) · 2025
The automatic generation of captions for satellite images can enhance the accessibility and interpretability of Earth Observation (EO) data. In this study, we compare two approaches to image captioning: TerraMind, a model developed within the FAST-EO project specifically for satellite imagery, and BLIP-2, a generic multimodal model trained on RGB images. The dataset used, SmallMinesDS, consists of annotated satellite images from five districts in Ghana, where unregulated small-scale gold mining threatens cocoa farmlands. Our evaluation focuses on caption accuracy, specificity, and adaptability to EO imagery, highlighting the strengths and limitations of each approach in the context of environmental monitoring.