Training data selection for voice conversion using speaker selection and vector field smoothing

M. Hashimoto, Norio Higuchi · 2002

We have previously proposed a spectral mapping method (SSVFS), for the purpose of voice conversion with a small amount of training data using speaker selection and vector held smoothing techniques. It has already been shown that SSVFS is effective for spectral mapping by both objective and subjective evaluations, and that it can operate with a very small amount of training data-as little as only one word (Hashimoto and Higuchi, 1995). We propose a criterion for selecting effective training data for SSVFS. We define coverage of parameter space with respect to the training procedure of SSVFS as the criterion. This criterion is useful not only for the selection of effective training samples, which is important for the efficient learning of spectral characteristics, but also for the estimation of the degree to which learning is carried out. To evaluate the validity of the proposed criterion, we measured the correlation between spectral resemblance and coverage. The result showed that the mean correlation coefficient for eight target speakers is -0.74 with the proposed criterion, and -0.59 without consideration of the training procedure. We conclude that the proposed criterion is useful in selecting effective training samples for SSVFS.

Read the paper · More papers on PaperTik