DDPMVC: Non-parallel any-to-many voice conversion using diffusion encoder

Ryuichi Hatakeyama, Kohei Okuda, Toru Nakashika · 2024

In this paper, we propose DDPMVC, a voice conversion (VC) model for non-parallel data that utilizes the diffusion model. The distinctive feature of diffusion models lies in their high expressive power for high-dimensional data and their ability to learn more stably compared to traditional generative models. Although several VC methods using diffusion models have been proposed, only VoiceGrad meets the non-parallel and any-to-many condition, which makes it easy to handle in training and inference. DDPMVC is regarded as an extension of VoiceGrad that incorporates a rule-based diffusion process into the encoder. By using the encoder to convert speech into latent variables that have less speaker information, it is expected to improve the accuracy of the non-parallel VC. Experimental results demonstrated that the performance of DDPMVC surpassed that of VoiceGrad in terms of the mel cepstral distortion and speaker similarity.

Read the paper · More papers on PaperTik