Mix-Guided VC: Any-to-many Voice Conversion by Combining ASR and TTS Bottleneck Features

Zeqing Zhao, Sifan Ma, Yan Jia, Jingyu Hou, Lin Yang, Junjie Wang · 2022 13th International Symposium on Chinese Spoken Language Processing (ISCSLP) · 2022

Due to the difficulty of obtaining parallel data, there are many works focus on non-parallel voice conversion(VC) recently. Bottleneck features(BNFs) from automatic speech recognition(ASR) and text-to-speech(TTS) models play an important role in feature disentangling for VC. In this work, we propose Mix-Guided VC, a non-parallel any-to-many voice conversion model by combining ASR-BNFs and TTS-BNFs. We demonstrate that ASR-BNFs and TTS-BNFs are complementary. ASR-BNFs are more robust especially in any-to-many tasks, but suffer from leaking source speaker’s timbre information; TTS-BNFs are closely correlated with text, but lack robustness. Experiments show that the proposed model achieves the best balance in speech quality, timbre similarity and robustness compares with baseline models. Furthermore, the whole modules in the proposed model can be trained jointly and no more pre-training data is needed.

Read the paper · More papers on PaperTik