Development of Phoneme Dominated Database for Limited Domain T-T-S in Hindi
Archana Balyan · International Journal of Artificial Intelligence & Applications · 2017
Maximum digital information is available to fewer people who can read or understand a particular language.The corpus is the basis for developing speech synthesis and recognition systems.In India, almost all speech research and development affiliations are developing their own speech corpora for Hindi language, which is the first language for more than 200 million people.The primary goal of this paper is to review the speech corpus created by various institutes and organizations so that the scientists and language technologists can recognize the crucial role of corpus development in the field of building ASR and TTS systems.This aim is to bring together all the information related to the recording, volume and quality of speech data in speech corpus to facilitate the work of researchers in the field of speech recognition and synthesis.This paper describes development of medium size database for Metro rail passenger information systems using HMM based technique in our organization for above application.Phoneme is chosen as basic speech unit of the database.The result shows that a medium size database consisting of 630 utterances with 12,614 words, 11572 tokens of phonemes covering 38 phonemes are generated in our database and it cover maximum possible phonetic context.