[Invited] Generative Model-Based Text-to-Speech Synthesis
Heiga Zen · 2018
Text-to-speech (TTS) aims to convert an arbitrary text into a corresponding speech waveform. This technology has progressed from a rule-based approach to a statistical one. The statistical approach includes concatenative, unit-selection TTS and model-based, generative TTS. Recent introduction of deep learning to generative TTS has drastically improved the naturalness of synthesized speech. It is also capable to modify speaker characteristics and adding emotions while preserving its high quality. This paper discusses the progress of the technology.