DO WE NEED VOICEDIUNVOICED CLASSIFICATION

F. Gales, Sabine Buchholz, Kate Knill, Yamato Ohtani, Masami Akamine · 2011

Most HMM-based TTS systems use a hard voiced/unvoiced classi­ fication to produce a discontinuous FO signal which is used for the generation of the source-excitatio n. When a mixed source excitation is used, this decision can be based on two different sources of infor­ mation: the state-specific MSD-prior of the FO models, and/or the frame-specific features generated by the aperiodicity model. This paper examines the meaning of these variables in the synthesis pro­ cess, their interaction, and how they affect the perceived quality of the generated speech The results of several perceptual experiments show that when using mixed excitation, subjects consistently prefer samples with very few or no false unvoiced errors, whereas a reduc­ tion in the rate of false voiced errors does not produce any perceptual improvement. This suggests that rather than using any form of hard voiced/unvoiced classification, e.g., the MSD-prior, it is better for synthesis to use a continuous FO signal and rely on the frame-level soft voiced/unvoiced decision of the aperiodicity model.

Read the paper · More papers on PaperTik