Utterance Alignment of Language Models for Effective User Simulation in Task-Oriented Dialogues
Xiang Luo, Jin Wang, Xuejie Zhang · IEEE Transactions on Audio Speech and Language Processing · 2025
Traditional user simulators often rely on manually designed agendas, resulting in generated responses lacking diversity and spontaneity. However, building user simulators with large language models (LLMs) heavily depends on selecting examples for in-context learning. Additionally, complex tasks and lengthy contexts can further burden LLMs, making generating responses that meet the desired user goals challenging. To tackle the abovementioned issue, this study introduces AlignUS, which combines a small language model (SLM) and another LLM to construct a user simulator for task-oriented dialogue systems. The SLM can produce basic dialogue actions and natural language utterances through training. Meanwhile, the LLM verifies the dialogue actions generated by the SLM and polishes the generated natural language utterances using its reasoning capabilities. Additionally, the natural language utterances polished by the LLM are aligned with the SLM, enhancing the SLM's ability to generate diverse utterances. Based on this, the burden on the LLM can be alleviated, as it only needs to handle a small portion of the task without case sampling. Extensive experiments conducted on the MultiWOZ dataset validated the effectiveness of the proposed approach, highlighting improvements in response quality and diversity. Compared to the current state-of-the-art (SOTA), our model demonstrates an improvement of 0.26 in success rate and completion rate for the MultiWOZ dataset over 100 dialogue turns. The code is accessible athttps://github.com/suntea233/AlignUS.