Navigation in virtual worlds via natural speech
Björn Wolfgang Schuller, Frank Althoff, Gregor McGlaun, Manfred K. Lang · OPUS (Augsburg University) · 2001
In this paper we propose a new approach enabling users to intuitively navigate in arbitrary virtual 3D-worlds via natural spontaneous speech. Three different approaches are considered and evaluated: First of all, two stochastic topdown decoders - a one- and a two-pass, split between acoustic and semantic layers. Besides these, a new two-pass decoder is being introduced. Two interfaces allow for multimodal integration: the different decoders can be used as homogenous competing instances already on the syntactic-semantic layer. A context sensitive intention decoder can dynamically constrain their recognition processes and translates their semantic structures into an abstract formal grammar. This concept enables heterogenous connection of additional input modalities on higher levels. The decoder can furthermore provide a measurement for the confidence even of a semantic interpretation based on acoustic confidences and by using rival decoders.