Spoken Language and Vision for Adaptive Human-Robot Cooperation

Peter Ford · 2007

Developments Language and MeaningCrangle and Suppes (1994) stated: "(1) the user should not have to become a programmer, or rely on a programmer, to alter the robot's behavior, and (2) the user should not have to learn specialized technical vocabularies to request action from a robot."Spoken language provides a very rich and direct means of communication between cooperating humans (Pickering & Garrod 2004).Language essentially provides a vector for the transmission of meaning between agents, and should thus be well adapted for allowing humans to transmit meaning to robots.This raises technical issues of how to extract meaning from language.Construction grammar (CxG) provides a linguistic formalism for achieving the required link from language to meaning (Goldberg 2003).Indeed, grammatical constructions define the direct mapping from sentences to meaning.Meaning of a sentence such as (1) is represented in a predicate-argument (PA) structure as in (2), based on generalized abstract structures as in (3).The power of these constructions is that they employ abstract argument "variables" that can take an open set of values.(1) John put the ball on the table.(2) Transport(John, Ball, Table ) (3) Event(Agent, Object, Recipient) We previously developed a system that generates PA representations (i.e.meanings) from video event sequences.When humans performed events and described what they were doing, the resulting input pairs allowed a separate learning system to acquire a set of grammatical constructions defining the sentences.The resulting system could describe new events and answer questions with the resulting set of learned grammatical constructions (Dominey & Boucher 2005).PA representations can be applied to commanding actions as well as describing them.Hence the CxG framework for mapping between sentences and their PA meaning can be applied to action commands as well.In either case, the richness of the language employed will be constrained by the richness of the perceptual and action PA representations of the target robot system.In the current research we examine how simple commands and grammatical constructions can be used for action command in HRI.Part of the challenge is to define an intermediate layer of language-commanded robot actions that are well adapted to a class of HRI cooperation tasks.This is similar to the language-based task analysis in Lauria et al. (2002).An essential part of the analysis we perform concerns examining a given task scenario and determining the set of action/command primitives that satisfy two requirements.1.They should allow a logical decomposition of the set of tasks into units that are neither too small (i.e.move a single joint) nor too large (perform the whole task).2. They should be of general utility so that different tasks can be performed with the same set of primitives. Spoken Language ProgrammingFor some tasks (e.g.navigating with a map) the sequence of human and robot actions required to achieve the task can be easily determined before beginning execution.Other tasks may require active exploration of the task space during execution in order to find a solution.In the first case, the user can tell the robot what to do before beginning execution, while in the second, instruction will take place during the actual execution of the task.

Read the paper · More papers on PaperTik