SP TELEPHONE SPEECH

Cristobal Corredor-Ardoy, Lori F Lamel, M. Adda-Deckel, Jodie Gauvain · 1998

In this paper we report on experiments with phone recognition of spontaneous telephone speech. Phone recognizers were trained and assessed on IDEAL., a multilingual corpus containing telephone speech in French, British English, German and Castillan Spanish. We investigated the influence of the training material composition (size and linguistic content) on the recognition performance using context-independent Hidden Markov Models and phonotactic bigram models. We found that when testing on spontaneous speech data, using only spontaneous speech training data gave the highest phone accuracies for tie four languages, even though this data comprises only 14% of the available training data. The use of contextdependent HMMs reduced the phone error across the 4 languages, with the average error reduced to 5 1.9% from the 57.4% obtained with CI models. We suggest a straightforward way of detecting non speech phenomena. The basic idea is to remove sequences of consonants between ~JNO silence labels from the recognized phone strings prior to scoring. This simple technique reduces the relative average phone error rate by 5.4%. The lowest phone error with CD models and Eltering was obtained for Spanish (39.1 %) with 4 language average being 49.1%.

Read the paper · More papers on PaperTik