HDRS: Hindi Dialogue Restaurant Search Corpus for Dialogue State Tracking in Task-Oriented Environment
Shrikant Malviya, Rohit Mishra, Santosh Kumar Barnwal, Uma Shanker Tiwary · IEEE/ACM Transactions on Audio Speech and Language Processing · 2021
Due to the rapid increase in the development of Task-oriented dialogue systems, the need for labelled dialogue corpus has become inevitable. For the Hindi language, there is no such dialogue corpus yet available. As a first attempt, we release a Hindi Dialogue Restaurant Search (HDRS) corpus and compare various state-of-the-art dialogue state tracking (DST) models on it. The corpus consists of 1.4 k human-to-human typed dialogues collected using Wizard-of-Oz paradigm. The paper starts with a brief description of the corpus by providing the details of features, corpus collection process and statistical analysis, then the performance of baseline NLU and DST models are investigated. Further, we experimented two categories of state-of-the-art belief state trackers: (1) Non-contextual pre-trained word embedding based DST models; (2) Contextual pre-trained BERT based DST models. All belief trackers follow a three-layered generic architecture. The category-1 models use the static domain ontology, while category-2 models have the capability to handle the dynamic ontology. The DST models are compared on joint-goal and turn-request accuracy. Global encoder and Slot-ATtentive decoders (GSAT) outperforms all the models with 83.25% joint-goal accuracy, followed by SUMBT.