Voice conversion based on deep neural networks for time-variant linear transformations
Gaku Kotani, Daisuke Saito, Nobuaki Minematsu · 2017
In voice conversion, deep neural networks are now being used as conversion models that map source features to target features. In this framework, it generally needs a larger amount of data to train more accurate conversion models. This condition, however, will reduce usability of voice conversion because a text-to-speech synthesizer can be built when a large amount of training data are available. We argue that we should take advantage of top-down knowledge that we have instead of preparing a large amount of data. This paper proposes a novel architecture using deep neural networks which can achieve superior performance of voice conversion. Our proposal is a network-based conversion that realizes only linear conversion but in even a time-variant way. Experiments show that naturalness improvement was observed in subjective assessments. It is considered that linear constraints at each time step prevent trained models from converting input features to unrealistic features.