Adaptive End-to-End Text-to-Speech Synthesis Based on Error Correction Feedback from Humans

Kazuki Fujii, Yuki Saito, Hiroshi Saruwatari · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

We propose an end-to-end text-to-speech (TTS) method that can intuitively correct accent errors in synthetic speech by using feedback from humans. State-of-the-art end-to-end TTS methods can synthesize high-quality speech, but humans can hardly interpret the black-boxed TTS system represented as a stack of neural networks. This reduced interpretability prevents humans from controlling the TTS system to correct errors in synthetic speech intuitively. In this paper, we focus on a method to involve human listeners in the process of accent error correction for synthetic speech generated by end-to-end TTS. Specifically, we build an end-to-end TTS model equipping a prosody predictor that estimates the change of pitch for each syllable from phoneme embeddings and context-aware word embedding. Then, we perform a human-in-the-loop (HITL) framework to correct errors of the prosody predictor using collective intelligence of human listeners. The results of Japanese TTS experiments show that our HITL framework can successfully correct accent errors and contribute to the quality of synthetic speech comparable to the conventional method requiring an accent dictionary for text analysis.

Read the paper · More papers on PaperTik