First steps towards text profiling for speech synthesis

Christina Tånnander, Jens Edlund · Digital Humanities in the Nordic and Baltic Countries Publications · 2019

We discuss an important yet under-studied domain of language and speech research: spoken text. Spoken text is language that was originally produced as text, then presented to recipients as speech. From a research perspective, this domain warrants special treatment, and we propose a classification that affords a structured approach based on a division of a linguistic message to be investigated into a primary (original) and secondary (studied) form. Secondly, we present the MTM Read Aloud corpus (MTM-RAC), a Swedish text and speech corpus built on in excess of 10,000 books. The corpus is closed access due to copyright restrictions on the material, but the methods developed and the results of our work on the corpus are available for use with similar corpora. MTM-RAC is designed with spoken text in mind and contains texts that have been read aloud in order to produce talking books, either by a human or using speech synthesis (i.e. text-to-speech) and the corresponding sound files. Finally, as the main purpose of the corpus is to explore and evaluate different aspects of text profiling for the purpose of reading aloud, we present first insights into this kind of profiling, based on experiments carried out on the corpus.

Read the paper · More papers on PaperTik