Performance comparison among HMM, DTW, and human abilities in terms of identifying stress patterns of word utterances
Nobuaki Minematsu, Yukiko Fujisawa, Seiichi Nakagawa · 2000
We have been focusing on applying speech technologies to pronunciation learning. In our previous study[1], a stressed syllable detector was implemented by using stressed syllable HMMs and unstressed ones. And using the detector internally, several systems were implemented[2]. However, their development did not necessarily require the use of HMMs as an acoustic modeling method. In this paper, an HMM-based method, a DTW-based method, and a human strategy only with visual inspection were compared in terms of their performance in judging whether two utterances of a word have the same stress pattern, e.g. récord and recórd. Here, one utterance was given by a Japanese learner and the other one was done by a native speaker. Experiments showed that HMMs gave us the higher performance