Rules for Synthesizing Prosodic Features of Speech—Preliminary Investigation
L. R. Rabiner, Harry N. Levitt, A. E. Rosenberg · The Journal of the Acoustical Society of America · 1968
In an earlier investigation [L. R. Rabiner, Bell System Tech. J. 47, 17–37 (1968)], a procedure for synthesizing speech by rule was described. The speech generated by this procedure had good intelligibility, but a machinelike quality. This was due, in part, to inadequate rules for controlling the prosodic features of the utterances (e.g., stressed vowel duration, fundamental frequency contour). In order to improve the quality of the synthesized speech, a preliminary investigation was carried out using three simple, declarative sentences. The stress pattern for each sentence was defined by assigning one of four possible stress levels to each vowel. Increments in duration and fundamental frequency were used as the correlates of stress, and various combinations of these factors were evaluated using a paired-comparison technique. One set of conditions was found to be highly favored for each of the three test sentences by a crew of six subjects. Using the results of this experiment, a set of rules was determined for controlling prosodic features in synthesizing simple, declarative sentences. Examples of synthetic speech demonstrating these rules will be played.