Embodied Language Grounding with Implicit 3D Visual Feature Representations

Mihir Prabhudesai, Hsiao-Yu Fish Tung, Syed Ashar Javed, Maximilian Sieb, Adam W. Harley, Katerina Fragkiadaki · 2019

Consider utterance the tomato to left of pot. Humans answer numerous questions about situation described, as well as reason through counterfactuals and alternatives, such as, is pot larger than tomato ?, can we move to a viewpoint from which tomato completely hidden behind pot ?, can we have an object that both to left of tomato and to right of pot ?, would tomato fit inside pot ?, and so on. Such reasoning capability remains elusive from current computational models of language understanding. To link language processing with spatial reasoning, we propose associating natural language utterances to a mental workspace of their meaning, encoded as 3-dimensional visual feature representations of world scenes they describe. We learn such 3-dimensional visual representations---we call them visual imaginations--- by predicting images a mobile agent sees while moving around in 3D world. The input image streams agent collects are unprojected into egomotion-stable 3D scene feature maps of scene, and projected from novel viewpoints to match observed RGB image views in an end-to-end differentiable manner. We then train modular neural models to generate such 3D feature representations given language utterances, to localize objects an utterance mentions in 3D feature representation inferred from an image, and to predict desired 3D object locations given a manipulation instruction. We empirically show proposed models outperform by a large margin existing 2D models in spatial reasoning, referential object detection and instruction following, and generalize better across camera viewpoints and object arrangements.

Read the paper · More papers on PaperTik