A High-Quality Melody-Aware Peking Opera Synthesizer Using Data Augmentation
Xun Yu Zhou, Wujin Sun, Xiaodong Shi · 2023
The performing art of Peking Opera places great demands on the singing skills of singers, including pronunciation, melody, role, personal style and emotional expression, which poses a great challenge to Peking Opera singing voice synthesis. In this paper, we propose OperaSinger, following the main architecture of FastSpeech 2, using the features from the musical score as input, while improving the encoder in FastSpeech 2 by employing a stack of melody-aware location-variable convolution blocks in parallel with feed-forward Transformer blocks to alleviate the lack of naturalness caused by ignoring relatively local features. Due to the limitation of publicly available opera data, we explore several novel data augmentation methods to boost the training of OperaSinger. Extensive experiment results have demonstrated that 1) OperaSinger can generate high-quality Peking Opera samples (MOS 3.80) with naturalness and expressiveness; 2) the proposed data augmentation methods effectively improve performance on both subjective and objective evaluations.