A Benchmark for Scene-Aware Referential Gesture Generation

Anna Deichler, Rishabh Dabral, Fethiye Irmak Doğan, Anindita Ghosh, Jonas Beskow · KTH Publication Database DiVA (KTH Royal Institute of Technology) · 2026

Referential gestures, pointing, indicating, and orienting the body toward objects in shared space, are fundamental to embodied com- munication. For virtual agents and physical robots operating in human environments, the ability to generate spatially grounded gestures is essential for disambiguation, instruction, and collabora- tive interaction. Yet, research on communicative gesture generation has largely focused on co-speech beat and iconic gestures, trained on corpora in which spatial grounding is absent or incidental. This lack of active research on referential gestures can be attributed to three key factors: datasets that pair gestures with 3D scene context are scarce, referential gesture generation lacks task formu- lation, and metrics for evaluating spatial grounding do not exist. In this work, we address all three gaps by introducing the MM- Conv Referential Gesture Generation Challenge. Specifically, the benchmark consists of three components: (i) a paired data release of 3,000 pointing-annotated clips from MM-Conv and SGS-HSI, with pointing-quality annotations and scene-disjoint splits; (ii) a task formulation that requires systems to produce spatially grounded reference gestures aligned with speech, without oracle apex timing or motion templates; (iii) a spatio-temporal evaluation protocol decomposing referential gesture quality into temporal alignment, spatial accuracy, and referent recall. We present a modular baseline based on OmniControl and position the benchmark as the founda- tion for the scene-aware gesture generation challenge at the 1st Workshop on Human–Scene Interaction at ECCV 2026. We envision this challenge as a testbed for the next generation of referential gesture synthesis works.

Read the paper · More papers on PaperTik