Speaker invariant characterizations of vowels, liquids, and glides using relative formant frequencies
Toby E. Skinner · The Journal of the Acoustical Society of America · 1977
Sperry Univac is developing a system for automatically recognizing words in conversational speech. The linguistically oriented procedure identifies sounds by their acoustic-phonetic correlates, forms an hypothesized sequence of sound segments for each phrase, and then matches lexical entries where similar sound segments occur. Problems to be overcome include the style of speech (exhibiting extreme reduction, coarticulation, and dynamic range), the restricted frequency range of the speech signal (telephone bandwidth), and the variety of speakers (including widely divergent males and females). To accomplish speaker independence in the acoustic-phonetic identification processes, relative (as opposed to absolute) formant frequency characterizations of vowels, liquids, and glides are being employed. While a particular sound has disparate formant frequencies when produced by different speakers, the relative relationships between the formant frequencies are nearly invariant. The nature of the relative formant frequency measures and the results achieved will be discussed.