A Novel Dysarthric Speech Synthesis system using Tacotron2 for specific and OOV words

Komal Bharti, Samiul Haque, Pradip K. Das · 2024

Among various speech disorders, dysarthria presents a unique challenge when it comes to end-to-end speech synthesis. Its severity and complexity further escalate if appropriate treatment is not taken. In this paper, we propose a speaker-adaptive dysarthric speech synthesis technique using the Tacotron2 model. We used this model to generate dysarthric speech utterances that already exist in the UASpeech database and for Out-Of-Vocabulary (OOV) words as well to expand the vocabulary size. By generating dysarthric speech from textual input, we achieved favorable Mean Opinion Score (MOS) ratings for both known and OOV words. This approach was successful in adapting to the inherent variability and intelligibility differences among speakers. To ensure the fidelity of our generated speech, we incorporated Dynamic Time Warping (DTW) to demonstrate the similarity between the original and generated waveforms. Additionally, we applied Waveform similarity-based Synchronized OverLap-Add (WSOLA) on the OOV words, ensuring a flawless evaluation process. This approach allows us to address the challenge of limited data availability by enhancing database and vocabulary size. This paper aims to generate synthetic dysarthric speech to expand the pool of available UASpeech database, facilitating advancements in automatic recognition of dysarthric speech. Audio samples for all generated words for 6 dysarthric speakers are available at https://github.com/samiulhaq4424/MOS

Read the paper · More papers on PaperTik