HILvoice:Human-in-the-Loop Style Selection for Elder-Facing Speech Synthesis
Xueyuan Chen, Qiaochu Huang, Xixin Wu, Zhiyong Wu, Helen M. L. Meng · 2022 13th International Symposium on Chinese Spoken Language Processing (ISCSLP) · 2022
Controllable speech synthesis has made great progresses over the last decades. State-of-the-art systems can provide flexible interfaces for configuring the styles of generated speech for target users. However, for a specific user group, e.g., the older adults, using the available configuration interfaces to select the styles that are favored by the group still needs to be investigated. Two main questions of such a style selection are (i) how to provide various options for the target users to pick; and (ii) how to effectively obtain the opinions from the target users. Since these two questions are highly correlated which makes it difficult to solve them separately, we propose a holistic framework to consider these two questions together by involving the target users in an iterative loop. We demonstrate by experimental results that the proposed framework can successfully select a speaking style preferred by the older adults than the default neutral setting. Analysis results show that the selected style has slower speaking rate, which coincides with previous studies on auditory perception of older adults.