Transcription bottleneck of speech corpus exploitation

Caren Brinckmann · Publication Server of the Institute for German Language (Institute for German Language) · 2008

While written corpora can be exploited without any linguistic annotations, speech corpora need at least a basic transcription to be o f any use for linguistic research.The basic annotation o f speech data usually consists o f time-aligned orthographic transcriptions.To answer phonetic or phonological research questions, phonetic transcriptions are needed as well.However, manual annotation is very time-consuming and requires considerable skill and near-native competence.Therefore it can take years o f speech corpus compilation and annotation before any analyses can be carried out.In this paper, approaches that address the transcription bottleneck o f speech corpus exploitation are presented and discussed, including crowdsourcing the orthographic transcription, automatic phonetic alignment, and query-driven annotation.Currently, query-driven annotation and automatic phonetic alignment are being combined and applied in two speech research projects at the Institut fü r Deutsche Sprache (IDS), whereas crowdsourcing the orthographic transcription still awaits implementation.

Read the paper · More papers on PaperTik