Towards Coherent Co-speech Gesture Generation: Multimodal Approaches for Emotional, Generalized, and Interactive Modeling

Xingqun Qi · 2026

Modeling 3D human motion from multimodal signals is a fundamental problem in embodied AI, and co-speech gesture generation represents a key instance of this problem by linking speech perception with non-verbal motion synthesis. As humanoid robots move toward natural communication in shared physical spaces, generating coherent speech-driven 3D human gesture representations becomes an essential step toward retargeting expressive and physically plausible motions to robotic embodiments. This motivates the key scientific question of this thesis: To ensure the naturalness and physical rationality of gestures retargeted to humanoid robots, how can we model the corresponding 3D human representation sequences so that they are vivid in emotional expression, generalizable to diverse speech inputs, and applicable to interactive dialogue scenarios? Addressing this question requires overcoming three major challenges. First, human emotions evolve continuously during communication, yet existing methods often ignore affective dynamics or assume a single static emotion label. Second, current gesture generation models are usually trained on limited and constrained datasets, which restricts their ability to generalize to diverse, in-the-wild speech. Third, most prior work focuses on solitary speaker settings, while real-world dialogue involves concurrent and asymmetric interactions between speakers and listeners. These challenges are further intensified by the scarcity of large-scale, high-quality 3D datasets that capture emotional transitions, broad speech variation, and interactive multi-person motion. This thesis provides a systematic study of coherent co-speech gesture generation by addressing three interconnected dimensions: emotional vividness, in-the-wild generality, and conversational interactivity. First, we introduce the task of speech-driven emotional transition gesture generation. To alleviate the lack of emotion-transition data, we construct high-fidelity transition speech datasets using large language model based transcript generation and audio inpainting. A weakly-supervised training framework is then proposed to model temporal associations between emotional gesture segments, enabling smooth and diverse 3D gesture transitions across affective states. Second, we present CoCoGesture, a framework for coherent 3D co-speech gesture generation from in-the-wild speech. By constructing the large-scale GES-X dataset and adopting a pretrain-then-finetune paradigm, CoCoGesture learns a broad gesture prior and adapts it to unseen speech inputs. The proposed Mixture-of-Gesture-Experts (MoGE) block further enables adaptive fusion between audio and gesture features, improving temporal coordination while pre-serving motion diversity. Third, we extend co-speech gesture generation to concurrent two-person conversations through Co3Gesture. We construct the GES-Inter dataset and introduce a bilateral cooperative diffusion framework with a Temporal Interaction Module (TIM) and mutual attention mechanism, enabling the model to capture dependencies between concurrent motions and synthesize socially coherent interactive gestures. Extensive experiments demonstrate the effectiveness of the proposed datasets and frameworks across emotional, generalized, and interactive co-speech gesture generation settings. Collectively, this thesis advances 3D human gesture representation modeling from isolated speech-motion synthesis toward more expressive, robust, and interaction-aware embodied motion generation. The proposed methods provide a foundation for more natural virtual avatar animation and offer an important step toward future humanoid robots capable of expressive and socially coherent non-verbal communication.

Read the paper · More papers on PaperTik