Multispeaker and Multilingual Zero Shot Voice Cloning and Voice Conversion
Ruby Dinakar, Anup Omkar, Kedarnath K Bhat, M. Sai Nikitha, P Aftab Hussain · 2023
This research study presents an integrated model for voice cloning and voice conversion, trained on English and French using the Open SLR Multilingual LibreSpeech Datasets. Voice cloning involves synthesizing the voice of a desired person from text, while voice conversion modifies the voice of a source speaker to resemble a target speaker. The proposed model combines spectrum-prosody-cycleGAN for voice conversion and utilizes speaker encoders, decoders, attention mechanisms, and a WaveNet-based WaveGAN vocoder for voice cloning. Additionally, the study focuses on zero-shot voice cloning and voice conversion, demonstrating promising results without explicit training on the target person's voice. This work opens avenues for efficient and versatile speech modification applications.