Towards Zero-Shot Multi-Speaker Multi-Accent Text-to-Speech Synthesis

Mingyang Zhang, Xuehao Zhou, Zhizheng Wu, Haizhou Li · IEEE Signal Processing Letters · 2023

This paper presents a framework towards multi-accent neural text-to-speech synthesis for zero-shot multi-speaker, which employs an encoder-decoder architecture and an accent classifier to control the pronunciation variation from the encoder. The encoder and decoder are pre-trained on a large-scale multi-speaker corpus. The accent-informed encoder outputs are taken by the attention-based decoder to generate accented prosody. This framework allows for fine-tuning with limited training data from multiple accents, and is able to generate accented speech for unseen speakers. Both objective and subjective evaluations confirm the effectiveness of the proposed framework.

Read the paper · More papers on PaperTik