Formant-based synthesis of singing.
Sten Ternström, Johan Sundberg · 2007
www.speech.kth.se Rule-driven formant synthesis is a legacy technique that still has certain advantages over currently prevailing methods. The memory footprint is small and the flexibility is high. Using a modular, interactive synthesis engine, it is easy to test the perceptual effect of different source waveform and formant filter configurations. The rule system allows the investigation of how different styles and singer voices are represented in the low-level acoustic features, without changing the score. It remains difficult to achieve natural-sounding consonants and to integrate the higher abstraction levels of musical expression. Index Terms: formant synthesis, singing 1. Background Singing synthesis at KTH has its roots in the 1970’s, when Sundberg and Gauffin modified the text-to-speech systems developed by Carlson and Granström. An analogue singing synthesiser called MUSSE was built by Larsson in 1977 [1]. It included vibrato and other song-specific features, and could be played with a piano keyboard and joystick, or be remote-controlled by a minicomputer running a rule system. In the 1990’s, several digital implementations of MUSSE were made by Ternström and Berndtsson [2]. The synthesis model described here is a descendant of these, built with Aladdin, a commercial DSP tool that was another outcome of this work