Do You Know How Humans Sound? Exploring a Qualification Test Design for Crowdsourced Evaluation of Voice Synthesis Quality
Moe Yaegashi, Susumu Saito, Teppei Nakano, Tetsuji Ogawa · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022
This paper explores the effect of crowd worker filtering criteria on the crowdsourced subjective evaluation of synthesized voice. Currently, crowdsourcing is being used for the subjective evaluation of synthesized voice. In this case, it is important to remove workers who do not satisfy the requester's requirements, but effective worker filtering criteria remain unexplored. In this study, we focused on sound quality evaluation and explored effective worker filtering criteria for subjective evaluation of synthesized voice. To filter workers who can evaluate sound quality, we designed a task that compares pairs of synthesized voices with different amounts of distortion and selects less distorted voices. In this task, some pairs included obviously distorted voices to measure the degree of attention to distortion in the evaluation of each worker. The following three criteria for worker filtering were defined: whether the evaluation focused on the amount of distortion (Understanding of Requester's Intent), whether the evaluation was consistent (Response Consistency Rate), and whether the evaluation was made with confidence (Response Confidence). We conducted an experiment to explore the effects of these criteria on the sound quality evaluation results on Amazon Mechanical Turk. Our experimental results implied that the measures of Understanding of Requester's Intent and Response Confidence for pairs containing obviously distorted voices were effective in identifying workers who did not respond seriously and who did not understand the intent of the experiment.