Speaker identification: Effects of noise, telephone bandwidth, and word count on accuracy.
Dallasarra Cassie, Dunn Aericka, Shanna White, Al Yonovitz, Herbert Joe · The Journal of the Acoustical Society of America · 2010
It is of great interest by the legal system to identify individuals by their voice. Controversy in this area has continued for nearly five decades with states divided with regard to the merits of voice identification. The American Board of Recorded Evidence [ABRE (1999)] has established standards for the determination of identification or elimination of speakers. Digital spectrographic techniques, including formant tracking and finer descriptive measures of speech, are dramatic improvements and allow a test of the ABRE standards using improved technology embodying similar principles set forth in an aural and spectrographic method. Ten speakers recorded synthetic sentences at two different times. Words were distorted by mixing with noise or by telephone bandwidth reduction. All possible speaker pairs of “elimination” were presented to qualified listeners as well as equal probability of “identification.” For the comparison, subjects were presented spectrograms, formant tracks, fundamental frequency, and the ability to listen to single words. Confidence ratings and determinations (elimination or identification) for each comparison word were made with the addition of added words(<20). The results will be discussed with regard to correct classification based upon the number of words required and the reduction of correct classification based upon the distorted conditions.