Phoneme-level speaking rate variation on waveform generation using GAN-TTS

Mayuko Okamato, Sakriani Sakti, Satoshi Nakamura · 2019

The development of text-to-speech synthesis (TTS) systems continues to advance, and the naturalness of their generated speech has significantly improved. But most TTS systems now learn from data using a deep learning framework and generate the output at a monotonous speaking rate. In contrast humans vary their speaking rates and tend to slow down to emphasize words to distinguish elements of focus in an utterance.

Read the paper · More papers on PaperTik