Modeling the depth of grounding through a generative cognitive model in referential communication
Ryunosuke Baba, Junya Morita, Ryuichiro Higashinaka, Yugo Takeuchi · Computers in Human Behavior Artificial Humans · 2026
ABSTRACT To advance our understanding of referential communication and common ground formation, this study presents a novel generative cognitive model that integrates deep neural networks for visual perception, image generation, and language captioning. Using the Tangram Naming Task (TNT), we simulate the sender–receiver interaction with modular processes replicating holistic cognitive strategies. Through controlled simulation experiments, we discovered that language generation plays a more crucial role than visual perception in the success of TNT, while intermediate image generation enhances linguistic diversity, which is a key aspect of natural communication. Our results bridge cognitive modeling and large-scale generative models, demonstrating how internal cognitive dynamics can be visualized and quantitatively evaluated. This study contributes to the growing field of cognitive-inspired human–AI communication and offers a foundation for designing communication systems where internal understanding can be modeled, visualized, and aligned to foster trust and prevent miscommunication.