Resource development and experiments in automatic south african broadcast news transcription.
Herman Kamper, Febe de Wet, Thomas Hain, Thomas Niesler · 2012
We present a description of the development and evaluation of a first South African broadcast news transcription system. We describe a number of speech resources which have been collected in the resource-scarce South African environment for system development purposes: a 20 hour corpus of South African English (SAE) broadcast news; a 109M word corpus of South African newspaper text collected for language modelling purposes; and a 60k word SAE pronunciation dictionary. The development of our system is based on similar state-of-the-art broadcast news transcription systems. Our system uses cross-word triphone HMMs, MF-PLP features and persegment cepstral mean and per-bulletin cepstral variance normalisation. Our final system obtains a word error rate of 24.6%. We find that, for newsreader data, Indian and Black South African English accents are recognised more accurately than the speech by White English mother tongue speakers. However, for the spontaneous speech found in interviews and crossings to other locations, the latter accent is associated with the best results, although for this speech the error rates are high overall. Finally, we consider the recognition of MP3-compressed audio and show that performance only deteriorates at high compression levels.