Exploring Impact of Prioritizing Intra-Singer Acoustic Variations on Singer Embedding Extractor Construction for Singer Verification

Sayaka Toma, Tomoki Ariga, Yosuke Higuchil, Ichiju Hayasaka, Rie Shigyo, Tetsuji Ogawa · 2024

We explored the significance of prioritizing acoustic variations within singers while training a singer embedding ex-tractor for singer verification. Neural networks effective for speaker verification, like ECAPA-TDNN, may lead to in-creased false rejections as the number of identifiable speakers grows. Therefore, it is crucial to enhance the variation in training data per speaker. Singing voice, with its diverse ex-pressions and singing techniques, is believed to exhibit more intra-speaker variability compared to spoken voice. This study aimed to investigate the impact of intra-singer acous-tic variations on training singer embeddings. Experiments conducted with a self-constructed Japanese singing voice corpus revealed the following findings: 1) the importance of using multiple songs per singer, 2) the limited significance of increasing the total number of identifiable singers, and 3) the detrimental effects of discrepancies in singing technique between enrollment and verification data.

Read the paper · More papers on PaperTik