The CHJ corpus

John Fry · 2003

This chapter introduces the CallHome Japanese (CHJ) speech corpus, a collection of 120 recorded telephone conversations that serves as the basis of our annotated corpus. CallHome refers to a family of speech corpora that were assembled by the LDC beginning in 1995. The collection was sponsored by the Large Vocabulary Conversational Speech Recognition project of the US Department of Defense. The CHJ corpus includes transcripts of each conversation as well as digitized speech data. Each transcript covers a contiguous five- or ten-minute segment taken from a recorded conversation lasting up to 30 minutes. The transcribers made judgments as to the sex, age group, and regional accent or dialect of each of the 269 speakers. The CHJ transcripts are distributed as text files containing EUC-encoded Japanese characters. Each speaker turn consists of a continuous stretch of speech that begins and ends with a pause. Pause breaks, and hence turn boundaries, were determined by the judgment of the transcribers.

Read the paper · More papers on PaperTik