in Continuous Speech Recognition
James L. Hieronymus · 1989
Coarticulation alters the vowel formant characteristics in continuous speech. Studies of isolated monosyllables in the literature suggest that some phonemes cause more severe distortions than others. The largest changes are caused by lrl, Ill, lwl. Unstressed vowels are most affected. Previous studies by Holmes [6] and by us [14] indicate that these effects are even larger for continuous speech. Vowel recognition algorithms which do not take context into account in continuous speech normally achieve correct recognition of approximately 75 % for the three top choices from the recognizer, with the first choice approximately 60 %.. By developing methods which explicitly model the phonetic context, higher levels of performance can to be achieved. An ongoing study is being made of all 16 of the American English vowels using subsets of the DARPA acoustic-phonetic data base. Formants are obtained and normalized for each talkers formant range based on one sentence. The resulting formant tracks are smoothed using splines and sampled at 9 equally spaced points in time within vowel centered triphone regions. Triphones with semi-vowels in them are clustered separately. These formant values are k-means clustered using subsets of the sampled formant values. Then additional supervised training is done using other parameters including duration. The resulting clusters are used as a classifier using a modified Euclidean distance from the cluster centers. This results in approximately 80 % first choice vowel recognition at the outer edges of the vowel quadrilaterial. Stressed vowels were found to have spectra which statistically were no more stable than unstressed vowels. Other techniques for classification are being explored.