Making Full Use of Chinese Speech Corpora
Thomas Fang Zheng · 2004
It is well understood that the speech databases play a very important role for speech recognition. It is a dream for speech recognition researchers to create more useful databases with smaller efforts. To achieve this goal, the database should be well designed at first, and tools and more information should be provided so that the databases can be made full use of. This paper will illustrate the criteria according to which the Chinese speech databases will be created for different purposes. The way of transcription will also be discussed, which is the first thing to do after the data creation. Then examples on how to learn knowledge from the created database for other research purpose will be given. 1 Purpose of Speech Corpora The speech corpora play a very important role in speech and language processing, and this has been aware of by the speech community for years. To make full of speech corpora efficiently, people in the speech community worldwide have established consortiums, such as LDC, ELRA, and so on. The purpose of speech corpora can be illustrated in Table 1 (Kuwabara