Effects of speaker adaptive training on tensor-based arbitrary speaker conversion

Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose · 2012

This paper introduces speaker adaptive training techniques to tensor-based arbitrary speaker conversion. In voice conversion studies, realization of conversion from/to an arbitrary speaker’s voice is one of the important objectives. For this purpose, eigen-voice conversion (EVC), which is based on an eigenvoice Gaus-sian mixture model (EV-GMM), was proposed. Although the EVC can effectively construct the conversion model for arbi-trary target speakers using only a few utterances, increase of the utterances used to construct the conversion model does not always improve the conversion performance. This is because the EV-GMMmethod has an inherent problem in representation of GMM supervectors. We previously proposed tensor-based speaker space as a solution for this problem, and realized more flexible control of speaker characteristics. In this paper, to aim larger improvement of the performance of VC, speaker adaptive training and tensor-based speaker representation are integrated. The proposed method can construct the flexible and precise con-version model, and experimental results of one-to-many voice conversion demonstrate the effectiveness of the proposed ap-proach. Index Terms: voice conversion, Gaussian mixture model, eigenvoice, Tucker decomposition, speaker adaptive training

Read the paper · More papers on PaperTik