Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification
Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura · 2016
This paper presents a technique for transforming F0in a framework of statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC method converts voice timbre of singing voices of a source singer into that of a target singer without using vocoder-based waveform generation. Although this method achieves high sound quality of the converted singing voices, its use is limited to only intra-gender conversion without the need of F0transformation. To make it possible to also use the DIFFSVC method for cross-gender conversion, we propose a method to transform F0of an input singing voice for the DIFFSVC. The proposed method is also based on direct waveform modification using overlap-add process and filtering process. Results of subjective evaluations demonstrate that the proposed DIFFSVC method with F0transformation significantly improves sound quality of the converted singing voices while preserving the conversion accuracy of singer identity in the cross-gender conversion compared to the conventional SVC with vocoder.