Natural Prosody Generation in TTS for Marathi Speech Signal

Madhavi Repe, Suresh Damodar Shirbahadurkar, Smita A. Desai · 2010

A Text-To-Speech (TTS) synthesizer is a computer-based system that should be able to read any text aloud, whether it was directly introduced in the computer by an operator or scanned and submitted to an Optical Character Recognition (OCR) system. Systems that simply concatenate isolated words or parts of sentences, denoted as Voice Response Systems, are only applicable when a limited vocabulary is required (typically a few one hundreds of words), and when the sentences to be pronounced respect a very restricted structure, as is the case for the announcement of arrivals in train stations for instance. In the context of TTS synthesis, it is impossible (and luckily useless) to record and store all the words of the language. It is thus more suitable to define Text-To-Speech as the automatic production of speech, through a grapheme-to-phoneme transcription of the sentences to utter. In this paper, we implemented natural prosody generation in TTS for Marathi Speech Signal. Till now TTS for many languages is done like English, Maindrain, Telgu etc. Work on TTS for Marathi Language is also done but not with natural prosody effect. The synthesis technique used is Concatenation.

Read the paper · More papers on PaperTik