AdaStyleSpeech: A Fast Stylized Speech Synthesis Model Based on Adaptive Instance Normalization
Yuming Yang, Dongsheng Zou · 2024
Stylized speech synthesis transforms text into a specific style of speech guided by reference speech. Despite recent advancements in speech synthesis, challenges persist in this domain, including limitations in quality, speed, and similarity. To address these issues, we introduce AdaStyleSpeech, a novel model for stylized speech synthesis. This model can directly extract the style vector from reference speech. By combining textual information and style vectors, AdaStyleSpeech effectively transfers content into stylized speech using adaptive instance normalization. Additionally, we present AdaGANSpeech, a multi-style synthesis model based on stylistic mutual information and generative adversarial networks. Unlike AdaStyleSpeech, it works faster and can generate more diverse speech without the need for a reference. Experimental results demonstrate that AdaStyleSpeech attains remarkable outcomes in synthesis quality, making it a State-of-the-Art solution in stylized speech synthesis. AdaGANSpeech addresses the AdaStyleSpeech’s reliance on reference speech guidance during the generation phase and exhibits notable advantages in speech diversity, clarity, and synthesis speed.