Personalizing synthetic voices for people with progressive speech disorders: judging voice similarity
Sarah Creer, Stuart P. Cunningham, Phil Green, K. Fatema · 2009
Abstract Inbuildingpersonalized syntheticvoicesforpeople withspeechdisorders, the output should capture the individual’s vocal iden-tity. This paper reports a listener judgment experiment on thesimilarity of Hidden Markov Model based synthetic voices us-ing varying amounts of adaptation data to two non-impairedspeakers. We conclude that around 100 sentences of data isneeded to build a voice that retains the characteristics of thetarget speaker but using more data improves the voice. Experi-ments using Multi-Layer Perceptrons (MLPs) are conducted tofind which acoustic features contribute to the similarity judg-ments. Results show that mel-cepstral distortion and fraction ofvoicing agreement contribute most to replicating the similarityjudgment but the combination of all features is required for ac-curate prediction. Ongoing work applies the findings to voicebuilding for people with impaired speech. Index Terms : HMM-based speech synthesis, voice similaritymeasure, voice banking 1. Introduction