The organization of high-level visual cortex is aligned with visual rather than abstract linguistic information

Adva Shoham, Rotem Broday-Dvir, Rafael Malach, Galit Yovel · Journal of Vision · 2025

Recent studies showed that the response of high-level visual cortex to images can be predicted by their linguistic descriptions, suggesting an alignment between visual and linguistic information. We hypothesized that such alignment is limited to textual descriptions of the visual content of the image and does not extend to abstract descriptions. We distinguish between two types of linguistic descriptions of visual images: visual text, which describes the image’s purely visual content, and abstract text, which describes conceptual knowledge unrelated to the immediate visual attributes. Accordingly, we tested the hypothesis that visual text, but not abstract text, will predict the neural response to images in high-level visual cortex. To that purpose, we used visual and language deep learning algorithms to predict the iEEG responses in humans to images of familiar faces or places. We generated two types of textual descriptions for the images: visual text, describing the visual content of the image, and abstract text, based on their Wikipedia definitions. We then extracted the relational-structure representations from a large language model (GPT-2) for the text descriptions and from a deep neural network (VGG16) for the images. Using these visual and linguistic representations, we predicted the iEEG responses to the images. Our findings showed that neural responses in high-level visual cortex were similarly predicted by the visual representation of the images and linguistic representations of the visual text, but not by abstract text. Frontal-parietal electrodes showed a reverse pattern. These results are in line with recent findings showing that textual descriptions of the content of the image predict the response to images also in the macaque’s brain. These findings demonstrate that visual-language alignment in high-level visual cortex is limited to visually grounded language.

Read the paper · More papers on PaperTik