Zero-Shot Cross-Lingual Text-to-Speech With Style-Enhanced Normalization and Auditory Feedback Training Mechanism
Chung Tran, Chi Mai Luong, Sakriani Sakti · IEEE Transactions on Audio Speech and Language Processing · 2025
In an increasingly globalized and interconnected world, the ability to communicate in more than one language is a vital skill that can reduce language barriers and promote cultural interaction. However, mastering multiple languages requires a significant investment of time and effort. Here, zero-shot cross-lingual text-to-speech synthesis (TTS) offers benefits to augment human communication by producing high-quality speech in multiple languages while preserving the original speaker's vocal characteristics. However, building such a system presents several challenges, including ensuring high-quality synthesis and achieving similarity between the synthesized speaker and the reference speaker, especially when training a model for low-resource languages. In this study, we propose a novel technique known as Style-Enhanced Normalization TTS (STEN-TTS) to achieve two objectives: preserving synthesis quality while simultaneously enhancing the ability of zero-shot adaptation with just a few seconds of reference for the purpose of cross-lingual synthesis. The model itself can also be trained with low-resource data, but using data of only 10 or 20 minutes is a major challenge. To improve the quality of synthesized audio in low-resource languages, we propose a combination of STEN-TTS with different training methods, including unsupervised text encoding, knowledge distillation, and an auditory feedback mechanism. An experimental evaluation was conducted in five languages (English, Chinese, Indonesian, Japanese, and Vietnamese), considering high- and low-resource training data as well as seen and unseen speakers. The proposed approach has shown its effectiveness in a high-resource setting, achieving a remarkable similarity (SMOS) of 3.44$\pm$0.17 for cross-lingual conversion as well as verification scores of 93.4% and 80.5% for seen and unseen speakers, respectively. The results in a low-resource setting, measured by phoneme error rates, also indicate a substantial improvement, with enhancements of approximately 3-4% . In this case, the quality of speaker verification remains consistently high, achieving scores of 90.0% and 78.0% for seen and unseen speakers.