Detailed spectral analysis of a female voice
Dennis H. Klatt · The Journal of the Acoustical Society of America · 1986
Several thousand DFT magnitude spectra have been produced for a selected sample of speech from a single female speaker having a pleasant voice quality. The speaker sustained a number of different vowels while undertaking different laryngeal gestures including (1) slow 1-oct pitch glides up or down and (2) reiterant imitations of different sentences, using either [?V] of [hV] replacements for the pattern of stressed and unstressed syllables in each sentence, where [V] is one of six English vowels. An attempt has been made to quantify the effects on the harmonic spectrum of changes to fundamental frequency, syllable stress, and position of the syllable in the utterance. Also of interest is the detailed way that voicing is initiated and terminated in laryngealized versus breathy onsets and offsets. Previous analyses that used inverse filtering to study the glottal waveform seem to have missed several important aspects of the source spectrum and its change over time. Our analysis reveals the presence of considerable random breathiness noise at frequencies above 2 kHz over portions of many utterances. The strength of the fundamental component and of the various formant peaks relative to overall rms energy in the signal under all of the conditions outlined above have been quantified. There is variation in both the general tilt of the harmonic spectrum and the strength of the fundamental component depending largely on the (presumed) degree to which the larynx is spread/constricted. Locations of spectral zeros have been studied and related to the duration of the open part of glottal period. Synthesis has also been attempted using a modified version of the Klattalk synthesizer, which provides direct control over voicing source open period, source spectral tilt, breathiness noise, and degree of instability in the fundamental frequency contour (the latter parameter being particularly important for natural synthesis of a vowel sustained at constant pitch). Sentences spoken by replacing all of the syllables by [?V] were somewhat easier to synthesize—than sentences spoken by replacing all of the syllables by [hV]. The reason appears to be that in the latter materials, additional spectral peaks—presumably related to tracheal resonances—were sometimes observed in the vowel spectra. Except for this possibly important aspect of voice quality, the synthesis results are very encouraging; source control parameters and formant parameters are monotonic slowly varying continuous functions, suggesting that rules for the synthesis of more natural female voices might be formulated as a next step. [Work supported in part by NIH.]