Can Multimodal LLMs Reason About Stability? An Exploratory Study with Insights from the LLMs4PCG Challenge

Mury Fajar Dewantoro, Febri Abdullah, Yi Xia, Ibrahim Khan, Ruck Thawonmas, Wenwen Ouyang · 2025

This study investigates the extent to which multimodal large language models (MLLMs) demonstrate physical reasoning capabilities in dynamic, visually grounded environments. We evaluate whether using images as context in a prompt can enhance the ability of MLLMs to simulate and predict the outcomes of physical interactions. Using Science Birds, an Angry Birds-like physicsbased platform, we design a suite of tasks that probe core competencies in binary, comparative stability, and forward simulation using visual and textual inputs. Our evaluation shows that MLLMs can perform visual reasoning tasks with quantifiable accuracy, especially when making predictions based on image-rich input. However, their performance varies considerably depending on the specific model architecture and the type of input modality used. These findings highlight that incorporating both visual and textual data is crucial for accurate physical inference. Nonetheless, the evaluated MLLMs still have substantial limitations in structured visual reasoning tasks. By systematically analyzing the strengths and weaknesses of different models, our work provides practical guidance for advancing MLLM-based physical reasoning and supports the development of future benchmarks and competitions in this area, including contributions to the LLMs4PCG challenge. We make our source code and raw data available for future research.11https://anonymous.4open.science/r/cog2025-game-physics-eval/.

Read the paper · More papers on PaperTik