Performance Comparison of Machine Learning Models Trained on Manual vs ASR Transcriptions for Dialogue Act Annotation
Muhammad Usman Malik, Mukesh Barange, Julien Saunier, Alexandre Pauchet · 2018
Automatic dialogue act annotation of speech utterances is an important task in human-agent interaction in order to correctly interpret user utterances. Speech utterances can be transcribed manually or via Automatic Speech Recognizer (ASR). In this article, several Machine Learning models are trained on manual and ASR transcriptions of user utterances, using bag of words and n-grams feature generation approaches, and evaluated on ASR transcribed test set. Results show that models trained using ASR transcriptions perform better than algorithms trained on manual transcription. The impact of irregular distribution of dialogue acts on the accuracy of statistical models is also investigated, and a partial solution to this issue is shown using multimodal information as input.