Evaluation of Two Automatic Formant Extractors
James L. Flanagan · The Journal of the Acoustical Society of America · 1956
An evaluation of two electronic devices designed to automatically extract the formant frequencies from continuous speech is described. Both formant extractors yield three continuous dc output voltages whose magnitudes, as functions of time, represent the frequencies of the first three maxima in the short-time spectrum of the input speech. The two devices operate, however, upon different principles. One extractor periodically scans the speech spectrum to determine the frequencies of the maxima, while the other periodically examines appropriately restricted segments of the spectrum to determine the frequency of the maximum within each spectral segment. Evaluation tests are described for determining quantitatively the accuracy and reliability of both formant extractors. Specially constructed speech material is fed into the extractors, and their outputs are compared in a prescribed manner to spectrographic analyses of the input speech. The results of the evaluation show that the best formant extractor (the second) follows the first formant of the vowels in speech within ± 150 cps greater than 93% of the time; it follows the second formant within ± 200 cps greater than 91% of the time. The first and second formant outputs of the same extractor are simultaneously within the above tolerances greater than 85% of the time. An analysis of the test results according to the constituent sounds in the speech material shows that the nasal consonants and the extreme front and extreme back vowels cause the most formant errors. Examples of the typical operation of both formant extractors are presented.