StyleFormerGAN-VC:Improving Effect of few shot Cross-Lingual Voice Conversion Using VAE-StarGAN and Attention-AdaIN
Dengfeng Ke, Wenhan Yao, Ruixin Hu, Liangjie Huang, Qi Luo, Wentao Shu · 2022
Voice Conversion (VC) aims to transfer the speaker timbre while retaining the lexical content of the source speech and has attracted much attention lately. Although previous VC models have achieved good performance, unstability can not be avoided when it comes cross-lingual scenario. In this paper, we propose the StyleFormerGAN-VC to achieve better cross language speech conversion, where variational auto-encoder is introduced to model the feature distribution of the cross-lingual utterances and adversarial training is applied to elevate the speech quality. In addition, we combine the Attention mechanism and AdaIN to make our model more generalized to unseen speaker with long utterance. Experiments show that our model performs stably in the cross-lingual scenario and gains well MOS evaluation scores.