How Far ahead Can Model Predict Gesture Pose from Speech and Spoken Text?

Ryo Ishii, Akira Morikawa, Shin-ichiro Eitoku, Atsushi Fukayama, Takao Nakamura · 2023

We investigated how far into the future nonverbal behavior can be predicted from speech and speech text. Specifically, we build a model that generates future behaviors from speech and speech text information and evaluate the quality of the generated behaviors. This helps to clarify how far into the future behavior can be accurately predicted. Our experimental results show that in Gesture Pose Generation using speech and speech text, on the basis of the input speech and text, the nonverbal behavior up to at least 500 ms ahead can be predicted with objective evaluation values that are the same as those when no future prediction is made. This result shows a new possibility for Gesture Pose Generation using speech and speech text to predict the future up to at least 500 ms ahead with no performance degradation.

Read the paper · More papers on PaperTik