Speaker Timbre Supervision Method Based on Perceptual Loss for Voice Cloning
Jiaxin Li, Lianhai Zhang, Zeyu Qiu · 2023
Current voice cloning models are only trained by approximating the generated speech to match the reference speech, without directly supervising the speaker timbre information during training. This approach is prone to interference from other factors in speech, such as semantic and prosodic information, which hinders the voice cloning model's ability to learn speaker timbre information. To address this issue, this paper proposes a speaker timbre supervision training method based on perceptual loss. Specifically, we use a speaker recognition model highly correlated with speaker information to extract speaker perceptual features and calculate speaker perceptual loss. Furthermore, we input high-level perceptual features into a speaker discriminator for adversarial training, which guides the voice cloning model to learn speaker timbre information more effectively. The experimental results show that both speaker perceptual loss and speaker discriminators based on perceptual features improve the cloned timbre similarities of the voice cloning model, and combining both yields even better results.