Probabilistic integration of joint density model and speaker model for voice conversion
Daisuke Saito, Shinji Watanabe, Atsushi Nakamura, Nobuaki Minematsu · 2010
This paper describes a novel approach to voice conversion using both a joint density model and a speaker model. In voice con-version studies, approaches based on Gaussian Mixture Model (GMM) with probabilistic densities of joint vectors of a source and a target speakers are widely used to estimate a transfor-mation. However, for sufficient quality, they require a parallel corpus which contains plenty of utterances with the same lin-guistic content spoken by both the speakers. In addition, the joint density GMM methods often suffer from over-training ef-fects when the amount of training data is small. To compensate for these problems, we propose a novel approach to integrate the speaker GMM of the target with the joint density model using probabilistic formulation. The proposed method trains the joint density model with a few parallel utterances, and the speaker model with non-parallel data of the target, independently. It eases the burden on the source speaker. Experiments demon-strate the effectiveness of the proposed method, especially when the amount of the parallel corpus is small. Index Terms: voice conversion, joint density model, speaker model, probabilistic unification