Compact speech representations for speech synthesis
Willem Bastiaan Kleijn, David T. Talkin · 2004
We describe a method for obtaining a compact speech-waveform representation based on frame theory (Mermelstein, P., 1973). In contrast to earlier frame-theory based representations, the frame used for the representation is continuously adapted to the signal, facilitating an accurate description of both stationary regions and of rapid transitions with a relatively low rate of coefficients (few parameters per second). The representation allows an accurate and unambiguous decomposition into voiced and unvoiced components and it is particularly suitable for time scaling and pitch scaling.