Footprint reduction of Concatenative Text-To-Speech synthesizers using polynomial temporal decomposition

Tamar Shoham, D. Malah, Slava Shechtman · 2010

High quality low footprint Concatenative Text-To-Speech (CTTS) synthesizers provide a persistent challenge in the field of speech processing. The spectral parameters representing the short speech segments used in the concatenation process constitute a large portion of the required memory. In this paper we propose to use a vectorial form of Polynomial Temporal Decomposition combined with jointly optimal segmentation and polynomial order selection in order to reduce the storage required for the spectral amplitude parameters by 50%, while preserving the perceptual quality of the obtained synthesized speech.

Read the paper · More papers on PaperTik