Analysis of Subjective Evaluation of Al Speech Synthesis Emotional Expressiveness
Yihua Ao, Miaotong Yuan · 2024
AI speech synthesis is widely used in audio books, movie and TV dubbing, virtual human and other fields. Speech synthesis technology is no longer limited to simply “understandable”, the ability to express emotion has become a key factor to measure whether AI speech can horizontally expand the application’s actual scene and vertically enhance the application’s potential. In this article, we use the MUSHRA (Multi Stimulus test with Hidden Reference and Anchor) subjective evaluation experiment for the emotional expression ability of AI speech synthesis. The MUSHRA subjective experiment method is used to score different timbres in a series of evaluation dimensions, compare the emotional expression ability of different timbres, and explore the relationship between the evaluation dimensions. Our experimental results show that the difference in perception of voice timbre with and without professional training in vocal broadcasting is not obvious. Different application scenarios have different requirements for the emotional expression of timbre, and certain timbre needs to be selected according to the usage. For the emotional expression ability of AI speech synthesis, the most important influencing factor is the speech speed adaptation, followed by the timbre adaptation, naturalness and fluency. Therefore, in the use of scenarios with emotional accuracy requirements, those highly adaptive timbre should be preferred, and in the voice adjustment at the post production stage, the speed of speech should be adjusted appropriately under the premise of ensuring the naturalness and fluency, which can effectively improve the emotional expression ability of timbre. The results are expected to provide new insights for the related research in the fields of artificial intelligence, deep learning on human voice with emotion study, as well as the exploration and practice in the fields of phonetics and intelligent speech.