The NTUT Blizzard Challenge 2010 Entry
Yuan‐Fu Liao, Ming-Long Wu, Shao-He Lyu · 2010
This paper describes our HMM-based speech synthesis system (HTS) [1] submitted to Blizzard Challenge 2013 [2]. The focus of this entry is to build a TTS without using any provided information and speedup the training procedures by parallel processing. In this system, the input text is tagged by Stanford parser [3] and transformed into phone sequences by Flite’s letter to sound module [4]. Then all utterances are force-aligned using a phone recognizer trained using TIMIT [5] corpus. To consider the relationship between neighboring sentences, the linguistic features beyond sentence level are extracted including the (1) number and forward and backward positions of sentences in a paragraph and (2) punctuation marks (PMs) of current and surrounding sentences. Moreover, deterministic annealing expectation and maximization (DAEM) [6] and minimum generation error (MGE) [7] criterions are used to initialize and fine-tune the HTS models, respectively.