Benchmarking Geospatial Visual Reasoning over Street Map Images

Haiting Zhou, Zhou Yu · 2024

In this paper, we present a novel VQA benchmark SM-VQA, which is built upon street map images. Specifically, SM-VQA contains about 9.5K real-world street map images collected from the open geospatial database OpenStreetMap. Each image in SM-VQA is also associated with detailed geospatial annotations, enabling it to automatically generate up to 50K distinctive QA pairs of five types of geospatial reasoning abilities. The evaluation of the state-of-the-art open-source and commercial LMMs reflects the great challenge posed by SM-VQA. The cutting-edge Phi-3V and GPT-4o models merely achieve accuracies of 45.3% and 46.7% respectively.

Read the paper · More papers on PaperTik