Attention, not scale, drives human-AI alignment in multimodal language prediction

Viktor Nikolaus Kewenig, Andrew Lampinen, Samuel A. Nastase, Christopher Edwards, Quitterie D’Elascombe, Akilles Rechardt, Jeremy I Skipper, Gabriella Vigliocco · npj Artificial Intelligence · 2026

Humans routinely draw on visual context to predict upcoming words. To what extent current vision-language models produce comparable behaviour is unclear. Here, we placed five state-of-the-art pretrained systems side-by-side with 600 human participants in a web-based Visual-World Paradigm. On each of 100 six-second movie clips, models and participants received either text-only or synchronised video + text and judged how likely a specified target word was to appear next; human eye movements were tracked throughout. Adding visual context increased model–human alignment in predictability ratings across all architectures (average Δ r ≈ 0.18) with no impact of parameter size. When visual context was informative, transformer attention significantly increased alignment. Attention maps from two transformer models corresponded with human gaze, explaining up to 70% of the inter-participant variance when the scene contained informative cues. Notably, cross-modal attention reliably tracked anticipatory human fixations on semantic cues. These results suggest that current transformer-based vision-language models can approximate human behaviour exploiting visual context during language prediction—and that selective attention to informative cues, not sheer model scale, is the principal driver of this alignment.

Read the paper · More papers on PaperTik