Extraction of similar terms for unsupervised utterance categorisation in technical support automated agents
Amparo Albalate, D. Dimitrov · 2007
In this paper we address the unsupervised automated categorisation of spoken language utterances within the context of a technical support automated agent. In particular, we analyse the role of feature extraction in the design of more accurate classifiers. The utterance classification is performed based on a K-means clustering algorithm. We then propose a feature extraction method consisting in the automatic identification of semantically equivalent terms. Finally, the performance of the resulting categoriser, in terms of accuracy, is experimentally compared against the basic K-means without feature extraction.