Speech Generation and Modification in Concatenative Speech Synthesis

Ingmund Bjørkan · NORA - Norwegian Open Research Archives · 2010

The work presented in this thesis is related to speech generation and speech modification in unit selection synthesis.A major problem in unit selection synthesis systems is the large variability in the synthetic speech quality due to audible discontinuities.Although using an exhaustive search in a large speech unit database, audible discontinuities occasionally occur due to concatenating speech units from different speech contexts which do not fit together acoustically.The main focus in this thesis has been to alleviate the problem of audible discontinuities at concatenation points.The thesis consists of a theory part and three experiment chapters.The first experiment chapter concerns the selection of speech units from the speech unit database, in order to better avoid audible discontinuities at concatenation points.A listening test on detecting discontinuities in vowel joins is presented.A comparison of different objective spectral distance measures was then performed, using the ratings from the listening test as a reference.In addition to classic spectral distance measures, a correlation based distance measure was tested in this experiment.This distance measure was found to be very correlated to pitch mismatches, and not so promising for detecting spectral mismatches.The distance measures' correlation to human ratings were however relatively low in this test.In addition, such a test would be influenced by the specific test design and the synthesis system.Hence, too certain conclusions can not be drawn from this experiment.Join cost function design based on perceptual experiments is then discussed, and a probabilistic join cost model is proposed.Another approach to alleviate the problem of audible discontinuities is to apply modification of the speech signal by the use of signal processing.The strategy is to apply a speech model, and then smooth estimated speech parameter trajectories across the concatenation points.Finally, synthetic speech can be reconstructed by the speech model.In this thesis, the use of modification by a harmonic speech model has been tested for smoothing of pitch discontinuities and spectral mismatches.The use of speech modification gives an additional need for robust speech i First and foremost, I would like to thank my supervisor Torbjørn Svendsen for all his help during the work of this thesis.Through valuable comments, corrections, discussions, review, and advices, he has contributed significantly to the understanding and progress during the work on this thesis.I would also like to thank all the coworkers on the Fonema project for interesting discussions and collaboration.In particular, I would like to thank Ingunn Amdal and Dyre Meen.Ingunn for her help with the experiment on applying the work on speech modification to the Festival unit selection system and for helpful comments and corrections to this thesis.I will thank Dyre for his contribution to the work with the ESPRIT algorithm for voicing and pitch estimation, by both sharing ideas and some Python code.Dyre is also acknowledged for his work on the speech synthesis processor, which made it possible to have a complete speech synthesis system working in the Python programming language.I am in addition grateful to Snorre Farner for numerous interesting discussions on speech synthesis and speech modification.In particular, I will thank for his contribution to v PREFACE the perceptual experiment presented in Chapter 7, and for the discussions on speech modification.I also wish to thank all my friends and colleagues at the speech processing group of NTNU for a tremendously nice working environment and for memorable social events.I am also very grateful to everyone that have spared their time for attending the listening tests in this thesis and made these tests possible.Finally, I will thank my parents for all their support, and my dear Hilde for her love, support, and encouragement.Together with our little boy Johan they have provided essential encouragement for this project.

Read the paper · More papers on PaperTik