LLM-Guided Speech Processing for 3D Human Motion Generation
Muhammad Hasham Qazi, Sandesh Kumar, Farhan Khan · 2024
This paper proposes a novel approach for generating 3D human motion from speech input using Large Language Models (LLMs). The system processes spoken language to derive motion generation prompts, which are then used by a 3D human motion generation engine to generate the corresponding 3D animations. This method eliminates the need for predefined action sequence text prompts and instead offers a dynamic approach towards motion generation. By utilizing a GPT model, the system produces body motions aligned with speech. Results demonstrate the system’s capability to generate various types of motions in response to diverse speech inputs. Despite challenges in handling complex sentences, motion continuity, and real-time performance, this approach lays the groundwork for applications in education, virtual reality-based training, and gaming environments, where naturalistic agent interactions are essential. Future improvements include optimizing the system for real-time applications and addressing challenges related to complex sentences and motion continuity.