Chinese MULTEXT: Recordings for a Prosodic Corpus
Masahiko Komatsu · Sophia linguistica · 2009
This paper describes the Chinese version of MULTEXT. MULTEXT is a multilingual speech corpus developed for research on prosody. The Chinese version includes 40 different passages read by 10 speakers (most speakers read 15 passages). The mean duration of the passages is 20.0 s, and the total length of the recordings is approximately 58 min. The present corpus adds a tonal language to the existing speech resource, namely, English, French, Italian, German, Spanish, and Japanese. Expanding the speech data to languages of various prosodic types brings our attention to the modeling of F0 contours. The MOMEL algorithm was originally developed for intonation languages and this must be considered when applying it to languages of other types. The recordings contained in Chinese MULTEXT have already been used in prosody research. The corpus will contribute to research on prosodic types. It is scheduled to be distributed for research purposes.