Methods for optimal text selection

Jan P. H. van Santen, Adam L. Buchsbaum · 1997

Construction of both text-to-speech synthesis (TTS) and automatic speech recognition (ASR) systems involves usage of speech data bases. These data bases usually consist of read text, which means that one has significant control over the content of the data base. Here we address how one can take advantage of this control, by discussing a number of variants of "greedy" text selection methods and showing their application in a variety of examples. 1. INTRODUCTION Both automatic speech recognition (ASR) systems and text to speech (TTS) systems have components that are trained on text---typically read text. Surprisingly often, training text is selected without giving much thought to optimality of the selected text. For limited domain situations, it may very well suffice to select randomly a subset from the domain for training purposes. In many ASR applications, and certainly in most TTS applications, however, the domain is open. And, as discussed at length in [7], in open domain situation...

Read the paper · More papers on PaperTik