Effects of cue impoverishment on intelligibility and naturalness of synthesized velar stops
Bathsheba J. Malsheen, James T. Wright, Melanie Yue · The Journal of the Acoustical Society of America · 1986
In a previous test of two MITalk-based synthesizers [B. J. Malsheen, J. T. Wright, M. Yue, and M. Peet, J. Acoust. Soc. Am. Suppl. 1 79, S25 (1986)], we found that the intelligibility of velar stops degraded more for one system than the other in a simulated telephone bandwidth condition. We hypothesized that this degradation was due to missing secondary cues normally present in human velar productions. In order to test this hypothesis we examined short-time spectra of the synthesized velar bursts and compared them with those produced by a male human speaker. We found that the synthesized velars lacked a number of acoustic cues present in the human productions. Secondary high-frequency energy peaks for initial and final velars before and after nonfront vowels were missing, primary peaks for front vowels were lower in frequency than those apparent in the human spectra, and primary peaks for nonfront vowels were higher. Velar stops were then resynthesized on the basis of the human model spectra. Preliminary results show that the more fully cued velars sound less “chirpy” and more natural both in normal and telephone conditions, and that intelligibility has improved.