Many-to-many voice conversion using hidden Markov model-based speech recognition and synthesis
Yoshitaka Aizawa, Masaharu Kato, Tetsuo Kosaka · The Journal of the Acoustical Society of America · 2016
This paper describes a many-to-many voice conversion (VC) technique that does not require a parallel training set of source and target speakers. In a previous study, we already proposed many-to-one VC method that consists of decoding and synthesis parts, and it does not require a parallel training set. The basic idea of this system is that an input utterance is recognized utilizing the hidden Markov model (HMM) for speech recognition, and the recognized phoneme sequences are used as labels for speech synthesis. The aim of this work is to extend functionality from many-to-one to many-to-many VC. In particular, we focus on the VC of emotional speech. In order to achieve this, we utilize speaker adaptation techniques to adapt HMMs for speech synthesis. By using adaptation techniques, an arbitrary speaker’s voice can be produced. In this work, a combination of constrained structural maximum a posteriori linear regression and maximum a posteriori estimation is used for adaptation. In order to evaluate the proposed system, subjective speech intelligibility tests were conducted. In the experiments, the proposed adaptation with 10 utterances was compared with traditional parameter training with 450 utterances. The results showed that speaker adaptation could be carried out without performance degradation.