Analysis of phonetic transcriptions for Danish automatic speech recognition
Andreas Søeborg Kirkedal · 2015
Automatic speech recognition (ASR) relies on three resources: audio, orthographic transcrip-tions and a pronunciation dictionary. The dictionary or lexicon maps orthographic words to sequences of phones or phonemes that represent the pronunciation of the corresponding word. The quality of a speech recognition system depends heavily on the dictionary and the transcrip-tions therein. This paper presents an analysis of phonetic/phonemic features that are salient for current Danish ASR systems. This preliminary study consists of a series of experiments using an ASR system trained on the DK-PAROLE corpus. The analysis indicates that transcribing e.g. stress or vowel duration has a negative impact on performance. The best performance is obtained with coarse phonetic annotation and improves performance 1 % word error rate and 3.8 % sentence error rate.