Improvement of quality of voice conversion based on spectral differential filter using STRAIGHT-based Mel-cepstral coefficients

Koike Harunori, Takashi Nose, Takahiro Shinozaki, Akinori Ito · The Journal of the Acoustical Society of America · 2016

In our previous research, we proposed a technique for converting the speech individuality of speech of an arbitrary input speaker into a specified speaker’s one. This technique used the neural network (NN) for many-to-one mapping, which was trained from the pairs of multiple source speakers and a target speaker. Using the NN, we design a filter based on the difference of the spectra of the source and the target speech to directly convert the waveform of the input speaker into the target speaker’s one. An advantage of the proposed method is that the direct waveform conversion alleviates the quality degradation caused by the F0 extraction error. To improve this proposed method, we use mel-cepstral coefficients extracted by STRAIGHT, the high-quality tool of speech analysis and synthesis. The higher order components of the cepstral coefficients determine the detailed shape of the spectral envelope, but the prediction accuracy of the high-order coefficients might be worse than that of the lower-order ones. Thus, we investigated the effect of the order of cepstral components to find the condition that gave better conversion quality. Additionally, we exploited the dynamic features to reduce the discontinuity of the frame and improve the converted speech quality.

Read the paper · More papers on PaperTik