Filler Prediction Based on Bidirectional LSTM for Generation of Natural Response of Spoken Dialog
Yoshihiro Yamazaki, Yuya Chiba, Takashi Nose, Akinori Ito · 2020
Most of the conventional response generation models do not generate speech disfluencies including fillers, because they are trained from a written language corpus. It is necessary to insert fillers to written sentences for training a response generation model for the spoken language. In this paper, we proposed the filler prediction model based on bidirectional LSTM (BLSTM). This approach can consider a whole utterance and model both positions and kinds of fillers simultaneously. The experiments showed that the proposed method surpasses the conventional approach in terms of the prediction accuracy.