Towards Human-Like Dialogue: A Low-Latency, Transition-Aware Talking Avatar Generation Framework
Joonsung Lee, B.‐J. You, Do Hyung Kim, Taehwi Lee, Soeun Baek, Soyeon Park, Chang-Gun Lee · 2025
Conventional conversational AIs that rely only on text or voice lack the immersive, face-to-face interactivity of human dialogue. This paper proposes a talking avatar generation framework that addresses two critical challenges for humanlike communication: response latency and visual continuity. First, the paper introduces a chunk-based streaming generation approach that breaks the avatar’s response into 1-second video chunks delivered incrementally. This ensures the avatar begins responding after a nearly constant short delay (independent of the total response length), closely mimicking human response timing. Second, a transition-aware chunk generation strategy is presented that seamlessly aligns the avatar’s motion when switching from silent (idle) to speaking, eliminating the unnatural visual jump at the start of the response. Built on the state-of-the-art MuseTalk lip-sync model, the framework is optimized with TensorRT to meet timing constraints for human-like communication. Consequently, the avatar maintains a stable response latency regardless of speech length, compared to the linear increasing delays of conventional methods. In addition to that the avatar’s transitions from silence to speech are smooth, preserving the natural flow of conversation.