Construction of Automatic Speech Recognition Model that Recognizes Linguistic Information and Verbal/Non-verbal Phenomena

Nagito Shione, Yukoh Wakabayashi, Norihide Kitaoka · 2023

This study addresses an automatic speech recognition (ASR) model that simultaneously recognizes linguistic information and various verbal/non-verbal (VNV) phenomena: fillers, laughter, interrogative ascending tones, end of speech, pronunciation errors, word fragments, etc. Although ASR technology has rapidly advanced in recent years, most ASR systems can only recognize linguistic information and can still not VNV phenomena. Recognition of VNV phenomena is important because speakers often convey their intentions during spontaneous speech using these types of utterances. We achieve VNV recognition by annotating VNV phenomena tags on learning transcribed text when the ASR model is trained. Our experimental results revealed that the accuracy of speech recognition is improved by also recognizing VNV phenomena. The ASR model in which the VNV phenomena tags appeared before the linguistic information in the transcribed text showed the best performance when using data from the Corpus of Everyday Japanese Conversation.

Read the paper · More papers on PaperTik