Automatic Speech Recognition Using Linguistic and Verbal/Non-Verbal Information

Nagito Shione, Yukoh Wakabayashi, Norihide Kitaoka · 2023

This study proposes an automatic speech recognition (ASR) model that simultaneously recognizes linguistic information and nine types of verbal/non-verbal (VNV) phenomena (fillers, laughter, small voice, interrogative ascending tones, end of speech, pronunciation errors, word fragments, speech related to the flow of conversation and speech in dialects or foreign languages). Although ASR technology has rapidly advanced in recent years, most ASR systems are only able to recognize linguistic information, and are still unable to detect VNV phenomena. However, recognition of VNV phenomena is important because speakers often convey their intentions during spontaneous speech using these types of utterances. We achieve VNV recognition by applying VNV phenomena tags to the transcription of the speech used to train the ASR model. Our experimental results when using data from the Corpus of Everyday Japanese Conversation revealed that speech recognition accuracy is improved when VNV phenomena are also recognized. We also observed that the best recognition performance was achieved when the VNV phenomena tags appeared before the linguistic information in the speech transcription used for model training.

Read the paper · More papers on PaperTik