Recent improvements of the RWTH large vocabulary speech recognition system on spontaneous speech
Achim Sixtus, Sirko Molau, Stephan Kanthak, Ralf Schlüter, Hermann Ney · 2002
The paper presents recent improvements of the RWTH large vocabulary continuous speech recognition system (LVCSR). In particular, we report on the integration of across-word models into the first recognition pass, and describe better algorithms for fast vocal tract normalization (VTN). We focus both on improvements in word error rate and how to speed up the recognizer with only minimal loss of recognition accuracy. Implementation details and experimental results are given for the VerbMobil task, a German spontaneous speech corpus. The 25.0% word error rate (WER) of our within-word baseline system was reduced to 21.4% with VTN and across-word models. Decreasing the real-time factor (RTF) by up to 85% resulted in only a small degradation in recognition performance of 2% relative on average.