Emotional Speech Generator by using Generative Adversarial Networks
Takuya Asakura, Shunsuke Akama, Eri Sato-Shimokawara, Toru Yamaguchi, Shoji Yamamoto · 2019
In this paper, we propose an affective voice conversion method that can generate an emotional phonation from neutral speech by using cycle-consistent generative adversarial networks (CycleGAN). Our method uses the Mel-cepstral coefficients (MCEPs), which are extracted from speech signal as an input. Next, we apply the modified network model which is comprised two components with a generator and discriminator. In this generator network, the pairing structure with an encoder and decoder is used for an accurate and fast calculation in the learning process. Furthermore, we construct two types of encoder; the one equips a content encoder for the linguistic-information, and the other equips a domain encoder for the emotional-information. This separation is enable to reproduce the smooth speech with emotional information. Finally, we evaluate the emotion expression and sound quality of speeches by using the subjective evaluation for an accuracy of emotional change. As the result, although it is necessary to improve the deterioration of sound quality, our method has accomplished to convert the emotion of the speech.