Automatic target generation for vowels

Hisao M. Chang, James D. Miller · The Journal of the Acoustical Society of America · 1986

Research in phonetic recognition by computer is of interest because it may offer better performance than systems using larger-recognition units such as syllables or words. Among several techniques developed to model speech signals as phoneme sequences (e.g., hidden Markov models and other feature-based methods) is the phonetic-target model based on the auditory-perceptual theory of Miller [ASHA Reports 14 (1984)]. In it: (1) talker differences that mask acoustic phonetic correlates are minimized when incoming speech signals are represented as paths in a three-dimensional space, where X = log(P3/P2), Y = log(P1/R), and Z = log(P2/P1); and (2) phonetic targets defined in this space seem to be uniquely related to the acoustic patterns of phonelike elements. An automatic procedure generated three phonetic targets for the vowels /ɪ/, /ɛ/, and /ʌ/ from a set of training tokens in stop-vowel-stop format recorded from two male and two female talkers. When four new talkers were tested, a vowel recognition accuracy of about 98 % was attained. Additionally, a segmentation algorithm isolated the vocalic segments from the stop bursts with 100% accuracy. [NINCDS grant R01-NS21994-01.]

Read the paper · More papers on PaperTik