PROSODY MODELS FOR CONVERSATIONAL SPEECH RECOGNITION
Mari Ostendorf, Izhak Shafran, Rebecca Bates · 2003
This paper describes a formal model for incorporating prosody in the speech recognition process, both for improving word recognition directly and for jointly recognizing words and underlying structure. The model includes the possibility of using an intermediate symbolic representation as well as direct conditioning on acoustic correlates. Alternatives for feature extraction are described, together with implications for statistical modeling. Examples of prosody conditioning in spontaneous speech recognition include acoustic model clustering and dynamic pronunciation modeling. 1.