Speaker-independent recognition of vocalic segments
Alexander I. Rudnicky · The Journal of the Acoustical Society of America · 1984
Speaker variations produce substantial differences in vocalic spectra: Vowel templates generated from one speaker's voice will not accurately match another voice. It is possible, however, to impose transformations on the spectrum that factor out speaker differences. The current work presents two sets of experiments that examine such transformations. The first set of experiments used steady-state vowels (/i e a o u/); for ten male and ten female talkers, the results indicate that a log-ratio transformation that incorporates pitch and formants into a three-dimensional L space [log(F1/F0),log(F2/F1),log(F3/F2)] allows 93% classification accuracy. By comparison, spectral matching gives 44% accuracy. Extended vocalic segments (e.g., as in “away,” /əʏ/) have dynamically varying formant patterns. In L space, these patterns appear as tracks. The seeood set of experiments investigates a recognition technique that uses the encoded track shape as part of the representation. Speaker-independent performance on isolated words was better than that obtained through DTW matching using a spectral (reel scale) representation. All experiments were run using automatic pitch and formant trackers developed at C-MU.