Generating Prosody-Aligned Gestures via Residual Vector Quantization With Uniform Regularization and Activity Loss
Hyoungki Choi, Hyeongseok Choi, Jinbeum Jang, Joonki Paik · IEEE Access · 2026
Co-speech gesture generation (CSGG) aims to synthesize gestures that reflect both the semantic content and prosodic features of spoken language. While recent data-driven methods, such as diffusion-based models, have improved the realism of generated gestures, they often face challenges such as difficulties in deterministic alignment with speech prosody, high computational costs, and inference latency. Specifically, synchronizing gestures with fine-grained prosodic cues remains a challenge due to limitations in the structure of latent representations. To address this, we propose a two-stage framework consisting of a gesture embedding network and a gesture generator. The embedding network utilizes residual vector quantization (RVQ) with multiple codebooks to encode motion segments into discrete representations that preserve both macro-dynamics and micro-details of human gestures. To ensure diverse and efficient codebook usage, we introduce a uniform regularization (UR) objective that mitigates codebook collapse. We also design an activity loss (AL) that preserves the motion energy of the original gestures, enabling the model to respond adaptively to speech rhythm and intensity through deterministic mapping. Experimental results across several benchmark datasets indicate that our method achieves competitive or superior performance compared to representative state-of-the-art (SOTA) models, including diffusion-based approaches, in terms of Fréchet Gesture Distance (FGD) and motion diversity, while offering substantial improvements in inference efficiency. These findings validate the effectiveness and efficiency of our model in generating natural, expressive, and rhythm-aware co-speech gestures.