Explorations in speaker normalization

John J. Ohala, Yoko Hasegawa · The Journal of the Acoustical Society of America · 1987

Some preliminary efforts are reported at speaker normalization based on the assumption that listeners are tacitly aware (a) that they are listening to an instrument (the human vocal tract) which when uniform (unconstricted) should have resonances whose lowest frequencies are spaced according to the ratios 1:3:5:7, etc., (b) that the exact frequencies of these resonances depend on the length of the vocal tract, and (c) that what is relevant in resonances from a nonuniform tract is the ratio of these lowest resonances to their uniform value. There are probably several cues the listener could use to estimate the size of the speaker; for starters, we use the mean (long-term) F4. The reference resonances for F1r, F2r, F3r, i.e., those from the “uniform” tract, are 1/7, 3/7, and 5/7, respectively, of mean F4. The normalization consists in converting measured formant frequencies, F1m, etc., into the ratios of log F1m/log F1r. The results of applying this normalization to the entire voiced portion of vowels in a variety of CVC syllables is reported. [Work supported by a Sloan grant to Berkeley Cognitive Science Program.]

Read the paper · More papers on PaperTik