Statistics of short-term spectral characteristics of fluent speech
A. Maynard Engebretson · The Journal of the Acoustical Society of America · 1983
Spectral studies of speech sounds have essentially centered around carefully spoken isolated words and phrases. It is well known that the spectral characteristics of these sounds differ from sounds of normal fluent speech. Statistical properties of the spectral characteristics of fluent speech based on short-term, power-spectral-density (PSD) measures of the speech signal were studied. A computer program was written to automatically calculate the PSD of contiguous windows of speech data and to further calculate a variety of histograms of comparative measures of the PSDs. Using cross correlation as a comparative measure, three types of histograms were computed: type(1) differences between adjacent windows, type(2) duration of “steady-state” segments, and type (3) differences between average PSD estimates of “steady-state” segments. Type (1) and type (2) histograms approximate exponential functions. Type (3) histograms resemble uniform distributions before tailing off to zero at large values. Typical values for the mean of the “steady-state” duration and the average rate of these “steady-state” segments is 35 ms and 2.4 segments per second, respectively, which is indicative of syllables. Typical results for a variety of talkers will be presented. [Work supported in part by NS 03856.]