The construction of a Chinese interlanguage corpus

Bin Wu, Yanlu Xie, Lulu Lu, Chong Cao, Jinsong Zhang · 2016

A Chinese interlanguage corpus lays a foundation of studying speech production, such as the typical pronunciation errors, of non-native Chinese speakers. Traditional Chinese interlanguage corpus has difficulty in covering important phonetic types such tones, syllables with context. This paper presents a construction of an interlanguage corpus which contains 103 sentences covering 394 syllable types and 174 tri-tone types bounded by prosodic boundary using a modified least-to-most-ordered algorithm. The corpus uses about half the size of traditional interlanguage corpus the Conversational Chinese 301 while achieving better coverage and less uneven distribution of syllable type and tri-tone type. More than 80% of words of the interlanguage corpus can be found in the word list of HSK4; about 13% of words found HSK5 and HSK6; about 3% beyond the vocabulary of HSK6.

Read the paper · More papers on PaperTik