Analyzing and classifying Indonesian spontaneous and dictated speech
Cil Hardianto Satriawan, Dessi Puji Lestari · 2016
The accurate recognition of spontaneous speech is crucial in achieving practical speech recognition. Statistical-based recognition models typically employ a large amount of read or dictated speech for training, which often yields poor spontaneous recognition performance. Many approaches have been forwarded to improve performance, including model adaptation and model switching. In an effort to improve Indonesian language spontaneous recognition performance, we attempt to pinpoint the acoustic differences between spontaneous and dictated Indonesian speech. At the phoneme level, we find that there are differences in the distribution and pronunciation of several key phonemes associated with filled pauses. Across speakers, there is a consistent reduction in segment duration and segment energy, with a less marked spectral reduction. Using these differences as a starting point, we train a number of classifiers that can accurately identify spontaneous and read Indonenesian utterances at F1 scores consistently above 90%. We show that classification is achievable by considering segment features and feature differences between consecutive segments, or “delta” and “delta-delta segments”.