Speaker Identification and Vocal Variability

John A. Starkweather, Wm. A. Hargreaves · The Journal of the Acoustical Society of America · 1962

In the course of speech processing for the machine recognition of verbal content, it is desirable to minimize the effects of speaker differences. On the other hand, expressive differences between speakers and an individual's variability from one time to another are of psychological interest for a number of applied possibilities. Such measures would be particularly useful if accomplished under conditions of natural speech without artificial control of speech content. Twelve speakers, all young college women, were recorded 6 times on each of 3 experimental days. Measurement of these recordings was made with a real-time spectrum analyzer which integates the output of a bank of one-third-octave filters over successive 2-sec intervals. Linear discriminant functions were derived from the spectrum-analysis data, and a decision tree developed for individual speaker identification. Considerable success was obtained in the identification of new speech samples of 60 sec or less from the same speakers. An error matrix of 16 speech samples per subject showed a mean error of 10.4%, with 5 subjects identified correctly from all samples.

Read the paper · More papers on PaperTik