Corpus-based techniques in the AT&t nextgen synthesis system
Ann K. Syrdal, Colin W. Wightman, Alistair D. Conkie, Yannis Stylianou, Mark C. Beutnagel, Juergen Schroeter, Volker Strom, Ki-Seung Lee, Matthew J. Makashay · 2000
The AT&T text-to-speech (TTS) synthesis system has been used as a framework for experimenting with a perceptuallyguided data-driven approach t o s p e e c h s y n thesis, with primary focus on data-driven elements in the \back end".Statistical training techniques applied to a large corpus are used to make decisions about predicted speech e v ents and selected speech i n ventory units.Our recent a d v ances in automatic phonetic and prosodic labeling and a new faster harmonic plus noise model (HNM) and unit preselection implementations have signi cantly improved TTS quality and speeded up both development time and runtime.