Actor-Double-Critic: Incorporating Model-Based Critic for Task-Oriented Dialogue Systems
Yen-Chen Wu, Bo-Hsiang Tseng, Milica Gašić · 2020
In order to improve the sample-efficiency of deep reinforcement learning (DRL), we implemented imagination augmented agent (I2A) in spoken dialogue systems (SDS).Although I2A achieves a higher success rate than baselines by augmenting predicted future into a policy network, its complicated architecture introduces unwanted instability.In this work, we propose actor-double-critic (ADC) to improve the stability and overall performance of I2A.ADC simplifies the architecture of I2A to reduce excessive parameters and hyperparameters.More importantly, a separate model-based critic shares parameters between actions and makes back-propagation explicit.In our experiments on Cambridge Restaurant Booking task, ADC enhances success rates considerably and shows robustness to imperfect environment models.In addition, ADC exhibits the stability and sample-efficiency as significantly reducing the baseline standard deviation of success rates and reaching the 80% success rate with half training data.