VietSing: A High-quality Vietnamese Singing Voice Corpus
Minh Vu, Zhou Wei, Binit Bhattarai, Kah Kuan Teh, Tran Huy Dat · 2024
This paper introduces a comprehensive Vietnamese dataset designed for singing voice synthesis (SVS). While there are extensive datasets available for widely spoken languages such as English and Chinese, resources for less common languages like Vietnamese are still scarce. The dataset, VietSing, comprises high-quality audio recordings and corresponding phonetic annotations, meticulously curated to support the development and evaluation of SVS systems. Detailed phonetic transcriptions and alignment with musical scores are provided to facilitate precise modeling of Vietnamese phonetics and prosody in song. We outline the data collection process, annotation methodology, and the challenges faced in ensuring linguistic and musical accuracy. Initial experiments using popular SVS model demonstrate the potential of VietSing to enhance the naturalness and intelligibility of synthesized Vietnamese singing voices. The dataset is not publicly available due to licensing restrictions. Interested researchers can request access to the dataset by contacting the corresponding authors. Access will be granted subject to the approval of the appropriate licensing agreements. Some audio samples and the code for our baseline evaluation method can be found at this link1.