Some Results on Speaker Verification Using Amplitude Spectra

George R. Doddington, Barbara M. Hydrick, B. Beek · The Journal of the Acoustical Society of America · 1974

Verification is based upon the similarity of six spectral/time matrices located within a test phrase with corresponding matrices defined during training. Each matrix is 0.1 sec long and is precisely located by scanning the test phrase for a best match with the reference matrix. Evaluation was performed over a data set including 50 “known” speakers and 70 “casual impostors.” Twenty percent female speakers were used. Five different phrases (including “We were away a year ago”) were collected in each session. Known speakers gave 100 sessions; impostors, 20. Data collection spanned 3.5 months. The first 50 sessions of each known speaker's data were used for training, the last 50 for test; 0.6% of the phrases yielded unusable data. When this happened, a substitute phrase from that session was used (two substitutions allowed, maximum). All impostor acceptance rates were determined for 2% true speaker rejection. A single fixed threshold was used for all speakers. Impostor acceptance rates were 2.5% for one phrase, 0.25% for two phrases, and 0.08% for three phrases. Five percent of known speaker data was labeled by the speakers as “not normal” because of respiratory ailments, etc. This data yielded a 4.5% reject rate for one phrase. Two professional mimics were employed to attempt to defeat the system. Each chose the five subjects he thought he could most easily mimic. Interactive trials with immediate feedback were of no apparent aid. Successful impersonation of about 5.5% for one phrase was achieved. No successful attempts for three phrases could be constructed from the mimic data. Reject rate for known speakers was plotted versus session number, at a nominal reject rate of 10%. Initial and final reject rates of 5% and 15%, respectively, indicate the necessity of adaptation in a practical system. [This work is supported by Rome Air Development Center.]

Read the paper · More papers on PaperTik