Exploring the Potential of Multimodal Large Language Models for Question Answering on Artworks

Alessio Ferrato, Carla Limongelli, Fabio Gasparetti, Giuseppe Sansonetti, Alessandro Micarelli · 2025

This paper investigates the application of a Multimodal Large Language Model to enhance visitor experiences in cultural heritage settings through Visual Question Answering (VQA) and Contextual Question Answering (CQA).We evaluate the zero-shot capabilities of LLaVA-7b (Large Language and Vision Assistant) on QA using the AQUA dataset.We assess how effectively it can answer questions about artwork, visual content, and contextual information through three experimental approaches.Our findings reveal that LLaVA demonstrates promising performance on visual questions, outperforming previous baselines but facing challenges with questions requiring contextual understanding.The selective knowledge integration approach showed the best overall performance, suggesting an efficient knowledge retrieval systems could enhance performance.Moreover, we show how to exploit such models to provide correct personalized answers using a well-established visitor model.

Read the paper · More papers on PaperTik