English Emotional Voice Conversion Using StarGAN Model
Ali Hamid Meftah, Adal A. Alashban, Yousef Ajami Alotaibi, Sid‐Ahmed Selouani · IEEE Access · 2023
The StarGANv2-VC model is a many-to-many non-parallel generative adversarial network (GAN) voice conversion (VC) model that has proven effective in style conversion tasks. This study aimed to investigate the scalability and diversity of the model for English emotional voice conversion (EVC) across different speakers and emotions. We carried out five experiments using an Emotional Speech Database (ESD), comprising a single speaker-multi-emotion experiment, a multi-speakers-multi-emotions experiment (gender-dependent), and a multi-speakers-multi-emotions experiment (gender-independent). We also assessed the effect of training set size and compared the performance of the StarGANv2-VC model with a CycleGAN model. Our study found that the StarGANv2-VC model accurately converted the pitch of the voice across all four emotions (neutral, happy, sad, and angry). However, the model’s efficiency in converting multi-emotions to multi-speakers was not as high as its efficiency in voice conversion for multi-speakers. Further research is needed in this area. We objectively assessed the quality of the converted speech using Mel-frequency cepstral distortion (MCD) and root-mean-square error (RMSE) for spectrum and prosody, respectively. Additionally, we conducted cross-emotion recognition using a convolutional recurrent neural network (CRNN).