Advancing Urban Development through Vision-Language Models: Applications and Challenges of Satellite Imagery Analysis

Nhat-Trinh Le, Nhat Tan Thai, Cao Vu Bui · 2024

This paper explores the integration of vision-language models, particularly focusing on the innovative CLIP (Contrastive Language-Image Pre-training) architecture developed by OpenAI. CLIP revolutionizes the way machines understand images by correlating them with textual descriptions, thereby broadening the scope of AI applications across various fields including remote sensing and urban planning. We specifically examine RemoteCLIP, a tailored adaptation of CLIP designed to enhance the analysis of satellite imagery for urban development purposes. Our study conducts a thorough evaluation of RemoteCLIP’s effectiveness in urban growth monitoring, infrastructure development, and environmental impact assessments, with a focused case study in Da Nang city. We explore its capabilities to integrate complex data sets, its precision in urban analytics, and address potential challenges such as data privacy, integration difficulties, and the accuracy of outputs. Through our detailed analysis, this study highlights the transformative potential of vision-language models in reshaping urban management practices, making a significant case for their broader adoption in smart city initiatives.

Read the paper · More papers on PaperTik