Six approaches to limited domain concatenative speech synthesis
Robert J. Utama, Ann K. Syrdal, Alistair D. Conkie · 2006
This paper (based on an MS Thesis by Robert Utama in the Electrical and Computer Engineering department at Rutgers University) describes 6 limited-domain Text-to-Speech (TTS) systems that are constrained to the digit string and natural number domains (cardinal numbers only). Each of the 6 unit selection-based concatenative TTS systems were implemented in MATLAB. We evaluate and discuss various factors that influenced the naturalness or overall quality of the synthesized voice. Some of the factors studied were the length and type of the synthesis unit and the extent of co-articulation represented in the recorded speech database. We show that it is possible to create a high quality limited domain TTS system either with maximal or with carefully controlled minimal effects of co-articulation.