Automatic Formant Extraction Using Linear Prediction

Stephanie McCandless · The Journal of the Acoustical Society of America · 1973

An algorithm is presented for automatically extracting the first three formant frequencies and amplitudes during voiced segments of continuous speech. It uses as input the linear prediction spectra and a segmentation parameter to indicate voicing and volume. A linear prediction spectrum is a smooth curve whose peaks correspond to pole positions, and therefore, in general, to formants. Complications arise, however, during the following situations: (a) Two formants may merge into one peak (as in r and w). (b) One formant may disappear, legitimately (as in n and m). (c) Spurious peaks may appear (as in nasalized vowels). (d) F4 may be present to confuse the issue. The algorithm chooses the middle of each high-volume voiced region as an anchor point, and branches out from there in both directions in time. At each frame, formant slots are filled with available peaks on the basis of frequency position relative to an educated guess. After some special routines to check for unused peaks and/or unfilled slots (including a formant enhancement program to separate merged peaks) the educated guess is updated with the formant frequencies decided in the current frame, to be used in evaluating the next frame. The algorithm has been implemented on the UNIVAC 1219 and the Fast Digital Processor, a fast special-purpose computer, and has proven successful with a large number of unrestricted sentences.

Read the paper · More papers on PaperTik