Arbitrary speaker conversion based on speaker space bases constructed by deep neural networks
Tetsuya Hashimoto, Daisuke Saito, Nobuaki Minematsu · 2016
This paper proposes a novel approach to construct a Deep Neural Network (DNN) based voice conversion (VC) system, where DNNs are integrated with speaker eigenspace. The proposed network consists of multiple DNNs and each of them converts input features to features corresponding to a base of eigenspace. Training of these DNNs is achieved with the assistance of Eigenvoice GMM (EVGMM). Experimental evaluations using one-to-many VC tasks show that the proposed method achieved better performance compared with that of EVGMM.