Towards automatic corpus preparation for a German broadcast news transcription system

Wolfgang Macherey, Hermann Ney · IEEE International Conference on Acoustics Speech and Signal Processing · 2002

When setting up a speech recognition system for a new domain, a lot of manual effort is spent on corpus preparation, i.e., data acquisition, cutting and segmentation of the audio material, generation of pronunciation lexica, as well as the definition of suitable training and test sets. In this paper we describe several methods that help to automate and thus to speed up this procedure. For this purpose, we assume that only a preliminary, partially incorrect textual transcription is available. The effectivity of the proposed methods is demonstrated with the development of a transcription system for the recognition of German broadcast news.

Read the paper · More papers on PaperTik