Corpus-Based Training of Action-Specific Language Models
Lars Schillingmann, Sven Wachsmuth, Britta Wrede · 2007
Especially in noisy environments like in human-robot interaction, visual information provides a strong cue facilitating a robust understanding of speech.In this paper, we consider the dynamic visual context of actions perceived by a camera.Based on an annotated multi-modal corpus of people who verbally explain tasks while they perform them, we present an automatic strategy for learning action-specific language models.The approach explicitly deals with the asynchrony of actions and verbal descriptions and includes an automatic parameter optimization based on a perplexity measure.Results show that a significant improvement of the word accuracy can be achieved using a dynamic switching of action-specific language models.