Emotional Speech Cloning using GANs
S. Sethu Selvi, Vignesh Anantharamakrishnan, Avaneesh Koushik, Sai K Akhil · 2021
Speech cloning is one of the most sought-after applications of deep learning. While there have been great strides in the field, many have been data-inefficient and have not yielded good results. Furthermore, most synthesized speech is monotonous. The challenge in trying to recreate a given individual's emotions with very limited data is another challenge by itself. In this paper, an alternate approach to speech cloning is proposed by exploring the possibility of treating synthesized voice and synthesized emotion as two separate entities and combining the outputs sequentially. The first part of the proposed network contains a neural voice synthesizer to generate non-emotional speech using as little data as possible. The output of this network is then combined with an array of different speaker emotions and passed to a modified version of CycleGAN network called EmoGAN with the aim being to seamlessly add in different emotions as required in the context of different sentences. The EmoGAN has been trained on two emotions from the Toronto Emotional Speech Set (TESS) database: sadness and anger. This is evident by the convergence plots of generator and discriminator and listening to the synthesized speech and evaluating the subjective quality.