Distributed Feature Representations for Out-of-domain Detection in Dialogue Systems

Rimpa Deka, Palash Pratim Dutta, Aparajita Dutta · 2023

In recent times, the increasing application of dialogue systems has drawn research interest toward improving the task of identifying out-of-domain (OOD) input in order to avoid catastrophic failures of the system. OOD inputs are utterances by the end user which do not lie within the scope of the current system. Several machine learning-based methods have been employed for this task. Classification-based and distance-based methods are widely used for this task. However, these approaches have limitations like choosing appropriate OOD samples and distance metric. We overcome the limitations by using a lightweight pipeline that applies a representation model to generate vector embeddings for in-domain (IND) and OOD samples in an unsupervised manner. We employ an LSTM model as a classifier trained only with the IND samples. This removes the challenge of curating a large collection of OOD samples representing the open-world scope of the OOD class. We use a dataset comprising fake OOD samples that are very similar to the IND samples. Our pipeline outperforms the CNN-based state-of-the-art model in the task of OOD detection, suggesting that the representation models can capture the features of IND and OOD distribution well. Furthermore, the LSTM classifier is more robust since it outperforms the baseline by 6% in terms of the AUROC score even when a more challenging set of fake OOD samples are used.

Read the paper · More papers on PaperTik