BISHOP|BLENDER: Spatially Grounded Language Understanding in 3D Modelling Software
Peter Gorniak · 2009
2004) we investigated how people describe objects in visual scenes using spatial language, and built a visually grounded language understanding system that performed well in understanding visually referring expressions. One application of this work is in applications that share a virtual world with the user, such as 3D modelling applications. Specifically, we here address the problem of selecting objects in a complex and cluttered 3D scene such as that shown in Figure 1. Using a 2D pointing device such as a mouse to select objects in a 3D scene is errorprone, because many objects at different distances from the viewer may share the same 2D screen space. Users currently have to solve this problem by manually rotating the scene to isolate the intended target in its own 2D area (as task that can be hard in a cluttered scene), or by using