Benchmarking Geospatial Visual Reasoning over Street Map Images
Haiting Zhou, Zhou Yu · 2024
In this paper, we present a novel VQA benchmark SM-VQA, which is built upon street map images. Specifically, SM-VQA contains about 9.5K real-world street map images collected from the open geospatial database OpenStreetMap. Each image in SM-VQA is also associated with detailed geospatial annotations, enabling it to automatically generate up to 50K distinctive QA pairs of five types of geospatial reasoning abilities. The evaluation of the state-of-the-art open-source and commercial LMMs reflects the great challenge posed by SM-VQA. The cutting-edge Phi-3V and GPT-4o models merely achieve accuracies of 45.3% and 46.7% respectively.