“Coffee So Good, You’ll Want to Slap Your Barista”: Evaluating AI Co-Creation Through Dialogic Actions

Oliver Bown, Kazjon S. Grace, Rodolfo Ocampo Blanco, Francisco Ibarrola · Creativity Research Journal · 2025

Recent years have seen the rise of large language models (LLMs) and multimodal generators (here referred to as “request-based” generative systems) that take natural language inputs. These suggest new dialogic forms of interaction for users of co-creative systems, in which the human and computer converse about creative goals, outputs, and concepts. Yet current request-based generative systems are poor at refined dialogic interaction, posing practical limits on how people can use these tools creatively. While there are rich bodies of literature studying the creative capabilities of AI systems, and their conversational and interactive capabilities and design, this paper specifically targets the role of dialogic interaction in human-machine co-creative interaction. We propose the Creative Dialogic Actions Framework for analyzing the efficacy of co-creative tasks in terms of creative-dialogic actions, looking at four subjects of dialogue – goals, roles, information about the world and information about personal stances – and then considering the intent, direction, and target of those creative dialogic actions. Our analysis of human-machine co-creative dialogs with LLMs considers how dialogue participants effectively co-create, aligning their understanding of goals, roles, information about the world and personal stances, by successfully inferring those action elements. We begin by surveying self-selecting creative users of LLMs to understand how important and how successful dialogic interaction is. The survey results suggest that dialogic interaction, while present, is of limited value to creative AI users in their pursuit of positive successful creative outcomes. We then apply a creative dialogic analysis to three creative tasks devised by the authors to highlight ways in which dialogic interaction succeeds and fails in interactions with OpenAI’s GPT-4. Our analysis highlights how GPT-4 is poor at inferring or directing dialogic actions about goals, roles, and personal stances, but better around information about the world. We propose that in light of these weaknesses, it is useful to think of a spectrum of dialogic interaction that includes pseudo-dialogic interaction, where the system creates the impression of dialogue but doesn’t meaningfully achieve dialogue, and weak asymmetrical dialogic interaction, where the system possesses some capability to achieve dialogic interaction, but largely leaning on the greater dialogic capabilities of the user. We consider how these limited dialogic capabilities relate to current LLMs’ narrow cognitive architecture, having no complex facilities to model an interlocutor’s expectations, goals, and understanding.

Read the paper · More papers on PaperTik