Low delay statistical singing voice conversion with direct waveform modification based on spectral differential considering global variance
Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura · The Journal of the Acoustical Society of America · 2016
This paper presents a low-delay statistical singing voice conversion (SVC) method with direct waveform modification based on spectral differential (DIFFSVC) considering global variance (GV). Statistical SVC based on a Gaussian mixture model (GMM) is a technique to convert singer identity of a source singer's voice into that of a target singer converting several acoustic features. We have successfully improved quality of converted singing voices by proposing the DIFFSVC considering the GV to avoid parameterization errors caused by vocoding process and alleviate over-smoothing effects. On the other hand, this method is based on batch-type conversion processing, and therefore, it is difficult to directly apply it to real-time conversion processing that needs a low delay conversion method. In this paper, to develop a high-quality and real-time SVC system, we propose a low-delay DIFFSVC method using GV-based post-filtering process. The experimental results have demonstrated that the proposed method makes it possible to improve speech quality of the converted singing voice in low-delay conversion processing. [This work was supported in part by JSPS KAKENHI Grant Number 26280060, Grant-in-Aid for JSPS Research Fellow Number 16J10726, and by the JST OngaCREST project.]