Decoding Visual Information from fMRI Data: A Multimodal Approach to Image and Caption Reconstruction
Matteo Ferrante, Tommaso Boccato, Furkan Özçelik, Rufin VanRullen, Nicola Toschi · Proceedings on CD-ROM - International Society for Magnetic Resonance in Medicine. Scientific Meeting and Exhibition/Proceedings of the International Society for Magnetic Resonance in Medicine, Scientific Meeting and Exhibition · 2024
Motivation: The study addresses the challenge of decoding and reconstructing visual experiences from fMRI data, an area yet to be mastered in neuroscience. Goal(s): We propose a methodology that deciphers brain activity patterns and renders these into visual and textual representations. Approach: We trained a linear model to map brain activity to image latent represenations. This informed a generative image-to-text transformer and a visual attribute-focused regression model, culminating in the creation of photorealistic images using a text-to-image diffusion model. Results: The model effectively combined high-level semantic understanding and low-level visual details, producing plausible reconstruction images from fMRI data. Impact: Our findings enhance our understanding of visual processing in the brain, with significant implications for integrating artificial intelligence (AI) with neuroscience.