Annotation and automatic recognition of spontaneously dictated medical records for Norwegian
V. Markhus, Bojana Gajić, J. Svarverud, L.E. Solbraa, Magne Hallstein Johnsen · 2004
In this paper we present a new research database of spontaneously dictated Norwegian speech, called MOBELspon, together with an experimental evaluation using standard automatic speech recognition (ASR) techniques. MOBELspon contains about 150 minutes of spontaneous dictation and 48 minutes of read speech of rheumatism health care records. The speakers are 10 medical students of both genders, coming from different parts of Norway and talking with their own dialect. MOBELspon contains a high degree of spontaneous speech features like disfluencies and para-linguistic speaker generated noise sounds. To model these features we propose some new special annotation symbols. MOBELspon contains the