Controllable Dialogue Simulation with In-context Learning
Zekun Li, Wenhu Chen, Shiyang Li, Hong Wang, Qian Jing, Xifeng Yan · 2022
Building dialogue systems requires a large corpus of annotated dialogues.Such datasets are usually created via crowdsourcing, which is expensive and time-consuming.In this paper, we propose DIALOGIC 1 , a novel dialogue simulation method based on large language model in-context learning to automate dataset creation.Seeded with a few annotated dialogues, DIA-LOGIC automatically selects in-context examples for demonstration and prompts GPT-3 to generate new dialogues and annotations in a controllable way.Our method can rapidly expand a small set of dialogue data with minimum or zero human involvement and parameter update and is thus much more cost-efficient and time-saving than crowdsourcing.Experimental results on the MultiWOZ dataset demonstrate that training a model on the simulated dialogues leads to even better performance than using the same amount of human-generated dialogues under the challenging low-resource settings, with as few as 85 dialogues as a seed.When enough data is available, our method can still serve as an effective data augmentation method.Human evaluation results also show that our simulated dialogues have near-human fluency and annotation accuracy.The code and data are available at https://github.com/Leezekun/dialogic. User([hotel]type is guest house, area is center): I need a place to stay.It should be a guest house in the center of town . GPT-3 GenerationTurn-level Belief State Dialog-level Belief State Dialogue Act System([hotel] [request] day people): how many days and people will be staying ?System([hotel] [offerbooked] reference [general] [reqmore]): the booking was successful.You reference number is [value_reference] .Can I help you with anything else ?User([hotel] stay is 2 , people is 3, day is Friday): we have 3 people, and we plan to stay 2 nights starting from Friday .