End-to-end Tibetan emotional speech synthesis based on Mandarin emotions transfer
Weizhao Zhang, Wenxuan Zhang · 2024
Emotional speech synthesis has attracted much attention in speech synthesis, especially in low-resource languages like Tibetan. However, Tibetan emotional speech synthesis is still in its infancy, facing challenges such as the lack of available public Tibetan emotional datasets and issues related to speaker disentanglement and emotional confusion. To address these problems, we propose an emotional Tibetan speech synthesis method based on improved FastSpeech2 training with a mix of neutral Tibetan, neutral Mandarin, and emotional Mandarin datasets. First, we replaced the normalization layer in the transformer structure of the original FastSpeech2 with an emotion-conditioned layer normalization(ELN), using emotion embeddings as conditional inputs to improve the model’s ability to learn emotions. Then, we used the orthogonal loss to disentangle the speaker vector and emotion vector to alleviate the speaker leakage problem. Additionally, orthogonal loss is also employed to address the problem of emotional confusion by enhancing the correlation between similar emotional features and ensuring independence among different emotional features. Experimental results showed that we successfully synthesized speech in five emotions for both Tibetan and Mandarin by transferring Mandarin emotions without any emotional Tibetan dataset. The mean opinion score (MOS) for all synthesized Tibetan speech was 3.87 or higher, while the MOS for all synthesized Mandarin speech was 3.81 or higher. Additionally, we evaluated the emotional transmission accuracy of the synthesized speech in Tibetan and Mandarin. The results showed that the emotional transmission accuracy of the Tibetan synthesized speech exceeded 80% in neutral, sad, and surprised emotions for the majority of speakers, while the emotional transmission accuracy of the synthesized Mandarin speech exceeded 69% across all five emotions.