Automatic Segment Alignment for Concatenative Speech Synthesis in Portuguese
Pedro Carvalho, Isabel M. Trancoso, Luís Oliveira · 2001
: Concatenative Text-To-Speech synthesizers join pre-recorded segments of speech data in order to produce high quality output speech. The synthesizer has to find the best segment to concatenate from an inventory of speech material. In order to do that, the inventory should be built from a correctly transcribed and time aligned speech database. This paper describes the construction of an automatically alignment tool using a Hidden Markov Model using very small training and test sets. Keywords: Alignment, Segmentation, Concatenative, Synthesis, Speech, HMM. 1. INTRODUCTION A generic text-to-speech system can be divided into two main modules: the first one performs text normalization and linguistic processing, and the second one generates the output speech waveform, using as input the string of phonetic symbols and prosodic parameters produced by the first module. Figure 1. A simple two module decomposition of a generic text-to-speech system The system that we are currently developing u...