Any-to-Any Voice Conversion with F0 and Timbre Disentanglement and Novel Timbre Conditioning
Sudheer Kovela, RAFAEL F. VALLE, Ambrish Dantrey, Bryan Catanzaro · 2023
Despite recent advances in voice conversion (VC), it is still challenging to do real-time one-shot voice conversion with good control over timbre and F0. In this work, we present a PPG-based VC model that directly decodes waveforms. We designed a speaker conditioned decoder based on HiFi-GAN[1], along with a new discriminator that produces high quality audio. Using an F0prenet and F0augmented speaker encoder, we are able to control F0and timbre independently with high fidelity. Our objective and subjective evaluations show that our method is preferred over others in terms of audio quality, timbre similarity and prosody retention.