Evolution and Perspectives of Speech Synthesis Technology: From Parametric Synthesis to the Era of Large Language Models

Yuhao Guo, Guanyu Li, Chenyu Xie, Qian Sun · 2025

Speech synthesis technology has evolved over six decades from rule-based to data-driven and now knowledge-driven approaches. This paper reviews its development, from early parametric methods (e.g., Formant, LPC) to modern Large Language Models (LLMs), highlighting key breakthroughs and paradigm shifts. Traditional methods relied on handcrafted acoustic rules, while deep learning (e.g., WaveNet, Tacotron) enabled end-to-end natural-sounding synthesis. Recently, LLMs like VALL-E and Voicebox have transformed speech generation into conditional language modeling, enhancing zero-shot cloning and cross-language synthesis. We propose a three-stage” division of this evolution and discuss three key features of speech synthesis in the LLM era: (1) joint text-to-speech modeling, (2) dynamic style control, and (3) adaptive generation with limited samples. Despite achieving near-human naturalness (MOS 4.5), current systems face challenges like high computational cost, insufficient emotional expressiveness, and security risks. Future trends include lightweight deployment, multimodal generation, and self-learning systems, with ethical guidelines needed to address deepfake risks.

Read the paper · More papers on PaperTik