Speaker Authentication Using a Formant-Tracking Vocoder

Jerry R. Richards · The Journal of the Acoustical Society of America · 1969

This talk describes some basic experiments centering around the use of linear discriminant functions to produce authentication judgments. A formant-tracking vocoder analyzer was used to reduce the speech waveform of key words to a set of eight time-varying parameters. Segmentations within the words were determined by noting rapid changes in the short-term power spectrum of each of the formants and changes in the state of voicing. Statistics derived from time-normalized segments within the words produced by a given speaker were used to generate an authentication mask for that speaker. The mask was automatically designed in a digital computer by utilizing Fisher's linear-discriminant function and some training utterances of the key words. After the masks were designed, they were used to classify new utterances and predict the likelihood that those utterances were uttered by a given speaker. The total system with the exception of the formant-tracking vocoder was computer simulated and testing was limited to a small group of five speakers. The system has the potentiality of being completely automatic in the learning and authentication phases, and the initial experiments were highly successful and indicate that reliable authentic results could be obtained using utterances from a set of five key words.

Read the paper · More papers on PaperTik