SleepWalker: Constrastive Fine-tuning Technique for Text to Kinematics Models for Human Computer Interaction
Demarcus Edwards, Danda B. Rawat · 2024
In this paper, we present SleepWalker, a few-shot fine-tuning approach for conditional text-to-motion generation. Our key innovation is extending contrastive learning, previously applied to image diffusion, to the new domain of text-to-motion synthesis. By binding subject representations to a text-conditioned diffusion model, SleepWalker enables personalized 3D human motion generation directly from natural language descriptions using just 15-25 examples. We demonstrate state-of-the-art performance on standardized metrics and benchmarks. The accessible fine-tuning method advances conditional motion synthesis, unlocking new creative possibilities in AR/VR, gaming, robotics, and human-computer interaction.