Knowledge-Driven Generative Design of Role-Playing Game Scenarios

Wojciech Owczarek, Julia Wróbel, Damian Pęszor · Applied Sciences · 2026

The paper addresses the problem of the generative creation of role-playing game scenarios based on a knowledge compendium. The purpose of this exploratory research was to determine the impact of elements of the generation process on the quality of the scenario. The research was conducted in a generative pipeline using large language models and a compendium that describes the world of the game. The scope of the study includes the creation of a modular system that enables ablation studies and the analysis of the influence of individual factors on the quality of the results. The experiments involved comparisons between language models, variants of knowledge compendia, and the count of user prompt steps. In addition, an ablation study, a self-bias study and a small-scale study with human respondents were conducted. The main purpose of these additional studies was to examine the methods used and identify potential problems regarding them. The ablation study supported the significance of creating a scenario skeleton in a non-random way. No indesputible self-bias was found. The human-based study showed that the LLM evaluators are, on average, less critical than the human ones, but share some similar scoring patterns. The study demonstrated statistically significant differences resulting from the choice of language model in Relevance, Coherence, Informativeness, Interactivity and Structure criteria, as well as the influence of the size of the compendium and the count of user prompt steps on the quality of the results. It was discovered that in the process of generating role-playing game scenarios, it might be beneficial to use short, non-randomly filled structures as the basis for the output scenario generation. It was found that large language models tend to score the generated scenarios higher than human respondents. There is, however, an overlap in preferences regarding the generation model between the human and the machine evaluators.

Read the paper · More papers on PaperTik