Learning of Object Properties, Spatial Relations, and Actions for Embodied Agents from Language and Vision

Muhannad Alomari, Paul Duckworth, David Hogg, Anthony G. Cohn · White Rose Research Online (University of Leeds, The University of Sheffield, University of York) · 2017

We present a system that enables embodied agents to learn about different components of the perceived world, such as object properties, spatial relations, and actions. The system learns a semantic representation and the linguistic description of such components by connecting two different sensory inputs: language and vision. The learning is achieved by mapping observed words to extracted visual features from video clips. We evaluate our approach against state-of-the-art supervised and unsupervised systems that each learn from a single modality, and we show that an improvement can be obtained by using both language and vision as inputs.

Read the paper · More papers on PaperTik