GlowVC: Mel-spectrogram space disentangling model for language-independent text-free voice conversion

Magdalena Proszewska, Grzegorz Beringer, Daniel Sáez-Trigueros, Thomas Wayne Merritt, Abdelhamid Ezzerg, Roberto Barra-Chicote · Interspeech 2022 · 2022

In this paper, we propose GlowVC: a multilingual multispeaker flow-based model for language-independent text-free voice conversion.We build on Glow-TTS, which provides an architecture that enables use of linguistic features during training without the necessity of using them for VC inference.We consider two versions of our model: GlowVC-conditional and GlowVC-explicit.GlowVC-conditional models the distribution of mel-spectrograms with speaker-conditioned flow and disentangles the mel-spectrogram space into content-and pitchrelevant dimensions, while GlowVC-explicit models the explicit distribution with unconditioned flow and disentangles said space into content-, pitch-and speaker-relevant dimensions.We evaluate our models in terms of intelligibility, speaker similarity and naturalness for intra-and cross-lingual conversion in seen and unseen languages.GlowVC models greatly outperform Au-toVC baseline in terms of intelligibility, while achieving just as high speaker similarity in intra-lingual VC, and slightly worse in the cross-lingual setting.Moreover, we demonstrate that GlowVC-explicit surpasses both GlowVC-conditional and Au-toVC in terms of naturalness.

Read the paper · More papers on PaperTik