MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching

Fabian David Schmidt, Florian Schneider, Chris Biemann, Goran Glavašš · 2025

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages.Consequently, evaluations of large vision-language models (LVLMs) predominantly target high-resource languages, underscoring the need for evaluation data for lowresource languages.To address this limitation, we introduce MVL-SIB, a massively multilingual vision-language benchmark that evaluates both cross-modal and text-only topical matching across 205 languages-over 100 more than the most multilingual existing VL benchmarks encompass.We then benchmark a range of open-weight LVLMs together with GPT-4o(mini) on MVL-SIB.Our results reveal that LVLMs struggle in cross-modal topic matching in lower-resource languages, performing no better than chance on languages like N'Koo.Our analysis further reveals that VL support in LVLMs declines disproportionately relative to textual support for lower-resource languages, as evidenced by comparison of cross-modal and text-only topical matching performance.We further observe that open-weight LVLMs do not benefit from representing a topic with more than one image, suggesting that these models are not yet fully effective at handling multiimage tasks.By correlating performance on MVL-SIB with other multilingual VL benchmarks, we highlight that MVL-SIB serves as a comprehensive probe of multilingual VL understanding in LVLMs. 1 * Equal contribution. 1 Code: https://github.com/floschne/mvl-sibk References 1 3 5 1 3 5 1 3 5 1 3 5 1 3 5 Images-To-Sentence: Select 1 of 4 sentences topically matching k reference images mSigLIP-base 57.7 64.6 66.4 53.3 58.6 59.7 51.4 56.2 57.1 38.9 41.2 41.7 36.1 37.6 38.0 Qwen2-VL 2B 36.3 34.8 34.9 35.5 35.3 34.1 34.5 34.3 33.2 31.0 30.6 30.0 29.5 29.0 28.6 Qwen2-VL 7B65.8 63.1 58.9 57.7 56.5 51.7 55.4 54.5 49.6 44.3 44.4 40.5 39.6 39.7 36.5 InternVL 2.5 4B 52.5 49.2 48.1 50.3 46.6 47.7 48.6 45.3 46.1 38.7 37.1 37.3 35.4 34.5 34.5 InternVL 2.5 8B 67.7 67.9 68.7 64.6 64.9 65.7 61.2 60.8 61.6 51.0 51.4 51.8 46.1 46.0 46.3 Centurio Qwen 54.8 60.0 62.4 54.2 59.2 60.6 53.4 58.1 58.9 46.6 48.9 49.2 43.0 44.2 44.7 GPT-4o-mini 68.3 78.1 77.4 71.6 79.0 78.1 72.0 78.9 77.7 63.5 68.0 66.4 56.9 60.3 58.7Sentences-To-Image: Select 1 of 4 images topically matching k reference sentences mSigLIP-base 56.3 66.0 69.6 51.8 61.6 64.0 49.1 58.3 60.2 36.0 40.4 41.2 32.9 36.3 36.9Qwen2-VL 2B 41.9 43.1 43.4 41.6 42.5 42.7 40.8 42.4 42.4 33.7 35.6 35.5 31.0 32.7 32.8 Qwen2-VL 7B 71.7 70.4 68.6 65.5 65.5 63.5 64.4 65.3 64.1 50.3 52.5 52.9 43.5 45.9 46.6 InternVL 2.5 4B 47.7 44.5 43.0 38.0 40.3 40.4 36.7 39.6 40.3 30.7 34.4 35.7 28.8 32.1 33.7 InternVL 2.5 8B 66.2 69.0 68.7 57.5 62.5 61.6 52.9 58.5 58.1 43.4 49.8 49.7 39.7 45.9 46.2 Centurio Qwen 35.3 36.1 35.6 31.1 32.9 33.3 31.0 32.8 33.1 28.7 29.7 29.8 28.1 28.7 28.7 GPT-4o-mini 77.5 86.4 89.1 77.2 86.5 88.6 77.1 86.1 88.4 68.4 79.8 82.7 61.7 74.0 77.2Drew A. Hudson and Christopher D. Manning.

Read the paper · More papers on PaperTik