A statistical vowel labeling technique
Wallace S. Tai · The Journal of the Acoustical Society of America · 1979
A vowel normalization technique based on the statistics of interrepetition variation by a single speaker has been implemented as part of the acoustic-phonetic analysis component of an SDC speech analysis/synthesis system. Since the statistics of interrepetition variation are represented by normal distribution [D. Broad, Phonetica 33, (1976)], a frequency-domain vowel normalization table can be constructed over each of the first three formant frequencies for the same speaker. The frequency range within which a vowel has the highest probability can be defined by calculating the intersections of the normal probability density functions of this vowel and its “adjacent” vowels. The probability for the vowel with the formant frequency falling outside the most probable range but within the normal distribution boundaries can then be obtained. Therefore, for each frequency interval, three or four vowel choices and their probabilities are included in the frequency-domain vowel table. For labeling an unknown vowel segment with steady-state formant frequencies, the candidate with the highest probability in the related formant interval is considered to be the best choice. Although more training sessions are required, an advantage of this labeling method over the traditional nearest-neighbor method where the linear distance between the formant frequencies of the unknown and the prestored targets are calculated, is its faster speed in a real-time environment because the number of computations is less. The study is also useful in testing the sensitivity to interrepetition variation of an analyzer which uses the nearest-neighbor method. Labeling results based on the statistical method and the nearest-neighbor method are compared for five male speakers. [Work supported by RADC.]