The deaccenting of given information in English in TTS systems: a case study

Alfonso Carlos Rodríguez Fernández-Peña · Archivum: Revista de la Facultad de Filosofía y Letras · 2024

This paper provides a descriptive qualitative and quantitative study of the deaccenting of given information, a.k.a. anaphora rule, by four well-known online TTS software (Murf, Lovo, Play.ht and Replica Studios). We have used 10 lines as input, each containing elements of given information to test the software. The voice types selected for our analysis are one male with British English accent and one female with American English accent for each software. Each line has been uttered by the voice skins in each software, downloaded in audio format and analysed using the speech analysis software Praat. This way we can measure and evaluate the pitch contours for each utterance and check whether the anaphora rule is applied or not by the different TTS software. The general results show that almost 70% of the lines do not achieve the delivery of the anaphora rule. This means that this prosodic feature characteristic of English stress and the substantial pragmatic load it carries is lost most of the times. The results obtained indicate that despite the fact that synthetic voices may be successful at segmental level in terms of catenation and voice quality, the suprasegmentals and prosodic elements of human speech are not mastered by the machines yet.

Read the paper · More papers on PaperTik