Analysis of Fundamental Frequency Contours in Speech

Harry N. Levitt, L. R. Rabiner · The Journal of the Acoustical Society of America · 1971

Many factors contribute to the variability of an ensemble of measured fundamental frequency (Fo) contours. These include, among others, differences between speakers, differences between replications of the utterance by the same speaker, and errors of estimation. Unfortunately, the relative contribution of the different sources of variability vary during the course of an utterance. We discuss in detail a method for analyzing such data by subdividing the time axis into a series of time windows and approximating the Fo contour within each window by a set of orthogonal polynomials. An analysis of variance is then carried out on the coefficients of each polynomial expression. Consistent trends in coefficient values between contiguous time windows are then approximated by sets of orthogonal polynomials. The technique is not only a powerful method for evaluating quantitatively the effects of different sources of variability, including changes in the degree of variability during an utterance, but also provides a basis for comparing directly the effect of short-term smoothing of contours with that of approximating long-term trends. Both these effects may also be compared directly with those of averaging over replications or averaging over both subjects and replications. These effects can be rank ordered in terms of a least-squares criterion, thereby isolating those factors which have the greatest influence on contour shape. The technique can also be used to separate out within-subject and between-subject differences in contour shape. An analysis of two utterances repeated three times by each of three subjects showed that inter-replication differences were relatively small and that a good fit to the measured fundamental frequency contours (rms deviation ≈2 Hz) could be obtained by using an average contour for each subject. A single mean contour could be fitted to the entire utterance, but in this case the rms deviation was close to 4 Hz. There were also substantial differences in variability during the course of the utterance, the stressed portions being less variable than the nonstressed portions.

Read the paper · More papers on PaperTik