Speech Gesture Generation from Acoustic and Textual Information using LSTMs
Ausdang Thangthai, Kwanchiva Thangthai, Arnon Namsanit, Sumonmas Thatphithakkul, Sittipong Saychum · 2021
This paper presents a gesture generation system developed for the generation and evaluation of non-verbal behaviour for embodied agents (GENEA) challenge 2020. The GENEA challenge provides approximately 4 hours of speech corpus, and 3D full-body human motions called the Trinity speech gesture dataset The gesture generation system's input varies from acoustic speech only, text only, acoustic and text, while the expected output is a video of the corresponding 3D motions.The proposed system consists of 2 steps: pre-processing the data and gesture modeling. In the first step, we pre-process the audio and text and stack them together as an input vector. We then pre-process the output of the 3D motion using a representation learning method. We select only 15 joints in the upper part of the body and represent them with a vector learned from the Denoising Autoencoder Networks (DAEs) model. In the second step, we build the gesture model using the encoder-decoder bidirectional LSTM architecture. This model transforms the acoustic speech and textual information into the DAEs output representation. We then decode the DAE representation to generate the 3D motion output video.The system has been evaluated using the held-out dataset and has been compared among nine systems: five participants, two baseline systems, two natural motions (real video). Two subjective evaluation methods, which are human-likeness and appropriateness, have been measured via a crowdsourced system. Our proposed system gets a fair rating score on both subjective evaluation methods.