Identification of normalized steady-state vowels

H. Wakita · The Journal of the Acoustical Society of America · 1975

This study was taken up as a first step toward automatic acoustic-phonetic transcription for arbitrary speakers. Sound identification experiments were conducted with nine steady-state American vowels /i, ɪ, ɛ, æ, ʌ, a, u, ᴜ, ɜ/. Reference information for each vowel was determined from vowel utterances in the context “hVd” produced by 16 speakers (nine males and seven females). In this case, the area function for each vowel was computed by use of the linear prediction method and was normalized to a reference length of 17 cm. Based on the resonance frequencies of the area functions thus normalized, the probability density function for the first three resonance frequencies of each vowel category was computed under the assumption of a normal distribution. These distributions were then used for the identification of steady-state vowels produced by an additional ten speakers (five males and five females). The overall rate of correct vowel identification was 83%. A higher recognition rate is expected to result from improving the estimation accuracy of the formant frequencies and bandwidths.

Read the paper · More papers on PaperTik