Physical modeling vocal synthesis

David M. Howard, M. Jack Mullen, Damian Murphy · The Journal of the Acoustical Society of America · 2006

Physical modeling music synthesis produces results that are often described as being ‘‘organic’’ or ‘‘warm.’’ A two-dimensional waveguide mesh has been developed for vocal synthesis that models the adult male oral tract. The input is either the LF glottal source model or a user-provided waveform file. The mesh shape, which is based on MRI data of a human vocal tract, can be changed dynamically using an impedance well approach to allow sounds such as diphthongs to be synthesized without clicking. The impedance well approach enables the mesh shape to be varied without removing or adding elements, actions that cause audible clicking in the acoustic output. The system has been implemented as a real-time MIDI-controlled synthesizer, taking its inspiration from the von Kempelin speaking machine, and this will be demonstrated live as part of the presentation. The system is set up to allow a continuous glottal source to be applied that includes vibrato, and thus the real-time output it currently produces is close to being vocalized. It should be noted, though, that appropriate variation of the lip opening does produce a voiced bilabial plosive, demonstrating the potential for moving towards a full speech synthesis system in the future.

Read the paper · More papers on PaperTik