A Learning-based Co-Speech Gesture Generation System for Social Robots
Xiangqi Li, Christian Dondrup · 2024
Co-speech gestures enhance both human-human and human-robot interactions. This paper examines the efficacy of a data-driven approach for generating synchronised co-speech gestures in three social robots to improve social interactions. Building on a sequence-to-sequence model, which maps speech to gestures [21], this work uses the Talking With Hands 16.2M dataset [11] to generate natural gestures for face-to-face conversations. Additionally, we address synchronisation issues identified in the original study. The model’s generality is tested on three robots—NAO, Pepper, and ARI. Objective and subjective evaluations, confirm that a data-driven approach effectively generates synchronised co-speech gestures.