Speaker-independent Spectral Mapping for Speech-to-Singing Conversion

Xiaoxue Gao, Xiaohai Tian, Rohan Kumar Das, Yi Zhou, Haizhou Li · 2019

Speech-to-Singing (STS) conversion aims at converting one's reading speech into his/her singing vocal. The prior work was mainly focused on transforming the prosody of speech to singing, however, there exist prominent differences between the spectra of speech and singing, which need to be transformed as well. In this paper, we propose to make use of parallel multi-speaker speak-sing data to develop a speaker-independent spectral mapping model, which is conditioned on i-vector to generate target speaker/singer identity. The model is therefore called speaker conditioned spectral mapping model. The converted singing spectra are then used together with prosody features to synthesize the target singing. We investigate the effectiveness of i-vector based average model adaptation to model the differences between speech and singing spectra for a specific speaker. The proposed model does not require parallel speak-sing data from target speakers during training. The experimental results conducted on NUS-48E and NUS-HLT-SLS database indicate that the proposed approach significantly outperforms the baselines in terms of quality and similarity.

Read the paper · More papers on PaperTik