System fusion for high-performance voice conversion
Xiaohai Tian, Zhizheng Wu, Siu Wa Lee, Quy Hy Nguyen, Minghui Dong, Eng Siong Chng · 2015
Recently, a number of voice conversion methods have been de-veloped. These methods attempt to improve conversion perfor-mance by using diverse mapping techniques in various acous-tic domains, e.g. high-resolution spectra and low-resolution Mel-cepstral coefficients. Each individual method has its own pros and cons. In this paper, we introduce a system fusion framework, which leverages and synergizes the merits of these state-of-the-art and even potential future conversion methods. For instance, methods delivering high speech quality are fused with methods capturing speaker characteristics, bringing an-other level of performance gain. To examine the feasibility of the proposed framework, we select two state-of-the-art meth-ods, Gaussian mixture model and frequency warping based sys-tems, as a case study. Experimental results reveal that the fusion system outperforms each individual method in both objective and subjective evaluation, and demonstrate the effectiveness of the proposed fusion framework. Index Terms: Voice conversion, system fusion, high-performance, frequency warping, GMM