Speech Recognition Simulation and its Application for Wizard-of-Oz Experiments
Alex Trutnev, Antoine Rozenknop, Martin Rajman · 2004
This contribution focusses on the simulation of speech recognition engines in the framework of Wizard-of-Oz experiments performed to design and evaluate dialogue-based vocal systems.Such simulation can be useful in cases where different dialogue management strategies are to be evaluated in terms of user satisfaction and where no automatic speech recognition engine or training data are available.The aim of the described work is to build a methodology and a tool that allow to simulate recognition errors in a controlled way.Two approaches are described, that produce, for any set of word sequences, simulated "recognition outputs" (i.e."noised" versions of the input sequences) in such a way that the obtained average Word Error Rate and Word Accuracy scores correspond as accurately as possible to pre-defined values representative of a targetted "true" speech recognition engine.The first approach aims at simulating Word Accuracy and Word Error Rates levels only.The simplicity of this approach leads to the fact that it does not fully simulate any real speech recognition engine, in the sense that it doesn't produce the same sentences as the real speech recogniser would.The second studied approach integrates the Viterbi decoding algorithm used in many speech recognition systems.The evaluation of the first approach was done using results produced by the Loquendo speech recognition system.The obtained average relative difference is of 1.54%.