Multimodal conversational systems for automobiles
Roberto Pieraccini, Krishna Dayanidhi, Jonathan E. Bloom, Jean-Gui Dahan, Michael Phillips, Bryan R. Goodman, K. Venkatesh Prasad · Communications of the ACM · 2004
The article reports that designing effective digital systems in safety-critical arenas takes interfaces included in the physical world. The multimodality of the system allows users to adapt to their environment, such as interacting through the graphic user interface (GUI) when the car is stopped at a light versus when the car is moving. In addition, the two modalities are designed to complement each other, the graphic interface controls providing hints of the corresponding voice commands. New users interact by speaking commands shown on the display, while the system engages in a directed dialogue and prompts for missing information. Experienced users can then adopt more effective commands and decide, at each turn, whether to interact by using speech or touch controls. The article also explains that the speech recognition engine makes use of dynamic semantic models that keep track of the current and past contextual information and dynamically modify the language model in order to increase the accuracy of the speech recognizer.