Tscan: Dialog Structure Discovery Using Scan, Adaptation of Scan to Text Data

Apurba Nath, Aayush Kubba · Engineering and Applied Sciences · 2021

Can we learn dialog structure from existing dialogs without ontology or domain assumptions. Understanding dialog structures from existing task oriented human human dialogs can help us automate these dialogues in a better way. Traditionally dialog structures have been created using ontologies that are created by domain experts. However, in our experience getting the ontology right is difficult and time consuming. Like other such tasks an unsupervised approach may do better than hand crafted rules. We propose an unsupervised dialog structure discovery approach that is based on SCAN (Semantic Clustering using Nearest Neighbors). Our approach comprises of two steps, the first being creating clusters of utterances and the second being creation of a structure using inter-cluster transition probabilities. Our main contribution in this paper is the adaptation of SCAN on text data. Unlike the SCAN approach for images, for text we did not train a separate pretext model and were able to use BERT for the same. Similarly for neigbor discovery, instead of augmentation we were able to leverage data variety. Evaluation metrics on dialog structures are a bit subjective, so we have used statistical measures as proxies for structure quality. We have also included our results on an internal human human task oriented 100k dialog dataset. We think SCAN like approaches are very promising for problems that use embedding similarities and should be further explored.

Read the paper · More papers on PaperTik