A new synthetic speech/sound control language

Osamu Mizuno, Shinya Nakajima · 1998

The Multi-layered Speech/Sound Synthesis Control Language (MSCL) proposed herein facilitates the syn-thesizing of several speech modes such as nuance, mental state and emotion, and allows speech to be synchronized to other media easily. MSCL is a multi-layered linguis-tic system and encompasses three layers: and seman-tic level layer (The S-layer), interpretation level layer (The I-layer), and parameter level layer (The P-layer). The S-layer is the description level of semantics such as emotional and emphasized speech. The I-layer is the description level of prosodic feature controls and inter-prets The S-layer scripts to for control on I-layer level. The P-layer represents prosodic parameters for speech synthesis. This multi-level description system is conve-nient for both laymen and professional users. MSCL also encompasses many eective prosodic feature con-trol functions such as a time-varying pattern descrip-tion function, absolute and relative control forms, and SDS(Speaker Dependent Scale). MSCL enables more emotional and expressive synthetic speech than conven-tional TTS systems. This paper describes these func-tions and the eective prosodic feature controls possible with MSCL. 1

Read the paper · More papers on PaperTik