Speech utterance categorisation given one training utterance per category
Amparo Albalate, David Suendermann · 2008
In this paper, we address the categorisation of speech utterances within the scenario of technical support automated agents given only one labelled utterance per category . The categorisation algorithm maps input utterances into bag-of-word vectors and then applies feature extraction based on soft word clustering. We analyse two feature extraction schemes: pole-based overlapping clustering (PoBOC) and a combination of PoBOC with Fuzzy c-medoids. For the categorisation at the utterance level, we use the Nearest Neighbour (NN) approach. Finally, we evaluate the proposed methods on a test corpus with more than 3000 utterances recorded in a commercial dialog system. (4 pages)