Speaker Authentication and Identification: A Comparison of Spectrographic and Auditory Presentations of Speech Material

Κ. Ν. Stevens, Carl E. Williams, Jaime R. Carbonell, Barbara Allen Woods · The Journal of the Acoustical Society of America · 1968

Speaker authentication and identification were examined for two different methods of presentation of the speech material: (1) speech samples presented aurally through headphones, and (2) speech samples presented visually as conventional intensity-frequency-time patterns, or spectrograms. Two kinds of experiments were carried out: (1) a series of closed tests in which there was a library of samples from eight speakers, and test utterances were known to be produced by one of the speakers; and (2) a series of open tests in which the same library of eight speakers was used, but test utterances may or may not have been produced by one of the speakers. The results for the closed tests indicate that, after about 4 h of exposure to the test situation, the percent error in identification of speakers from isolated speech samples (words or phrases) is about six percent for aural presentation and about 21% for visual presentation. These scores depend upon the talker, the subject, and the phonetic content and duration of the speech material. For the open visual tests, appreciable numbers of false acceptances (incorrect authentications) were made. The results suggest procedures that might be used to minimize error scores in practical situations.

Read the paper · More papers on PaperTik