Automatic estimation of extra-linguistic information in speech and its integration into recurrent neural network-based language models for speech recognition

Shohei Toyama, Daisuke Saito, Nobuaki Minematsu · The Journal of the Acoustical Society of America · 2016

When one talks with others, he/she often changes the way of lexical choice and the speaking style depending on various contextual factors. For example in Japanese, we often use gender-dependent expressions and, when we speak to elderly people, we use polite expressions. It can be said that in any language, formality or informality of speaking changes drastically depending on the situation where speakers are involved. The aim of automatic speech recognition (ASR) is to convert speech signals into a sequence of words and the above fact indicates that the performance of ASR can be improved by taking the various contextual factors into account in the ASR modules. In this study, at first, we attempt to estimate those contextual factors as para-linguistic or extra-linguistic information and then, to integrate the results into language models based on Recurrent Neural Network (RNN) for speech recognition. In experiments, from an input utterance, i-vector and openSMILE features were extracted to represent speaker identity and speaking style. These acoustically driven features were integrated into the reranking process of RNN-based language models. Reductions of perplexity of the language models were shown to be 1 to 2% relative.

Read the paper · More papers on PaperTik